Compare commits
885
Commits
6b30768853
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5fa7750046 | ||
|
|
1838eed855 | ||
|
|
96d8a5e3f0 | ||
|
|
0431e720de | ||
|
|
7f4921de49 | ||
|
|
d4c19baa32 | ||
|
|
d444fd8036 | ||
|
|
d42b1e2dda | ||
|
|
9492dc4c39 | ||
|
|
172beca3c5 | ||
|
|
f0c1289519 | ||
|
|
128172d3a8 | ||
|
|
820e8325a6 | ||
|
|
74c6a0f5eb | ||
|
|
d289a9101c | ||
|
|
9553ecb16a | ||
|
|
fb104bf05a | ||
|
|
f4adc31215 | ||
|
|
c6254f2342 | ||
|
|
d25b147a56 | ||
|
|
1f9915f074 | ||
|
|
c0e7a97a56 | ||
|
|
68c835f33b | ||
|
|
6b9fda76db | ||
|
|
1be259b66d | ||
|
|
2c3c88d067 | ||
|
|
3fc14d16bf | ||
|
|
4cc50889d6 | ||
|
|
63e68740b6 | ||
|
|
f2fddeb87d | ||
|
|
eda411c0be | ||
|
|
492757ce70 | ||
|
|
b58282b3d6 | ||
|
|
9da2760a42 | ||
|
|
049f633667 | ||
|
|
84946ed0c6 | ||
|
|
514e13660c | ||
|
|
9333334b7b | ||
|
|
367d0203b8 | ||
|
|
37017882c7 | ||
|
|
f116a1584e | ||
|
|
1a5bacce87 | ||
|
|
f6e5f7dcd6 | ||
|
|
0f0c407458 | ||
|
|
354992c9e7 | ||
|
|
28987240aa | ||
|
|
fbf2b92167 | ||
|
|
206a119a4b | ||
|
|
671e7ea5a4 | ||
|
|
d5cf3db2ec | ||
|
|
dc8823724d | ||
|
|
a392108562 | ||
|
|
fd1ab58598 | ||
|
|
bd13bdc6e1 | ||
|
|
bdf1bce4be | ||
|
|
c5c0477653 | ||
|
|
76faf0f9f8 | ||
|
|
879109e55d | ||
|
|
61b7d2a96a | ||
|
|
15d802ae3d | ||
|
|
ac17be2cd9 | ||
|
|
f1ba418aa8 | ||
|
|
b17cc3a09d | ||
|
|
1c38abe1d0 | ||
|
|
5d42a28d34 | ||
|
|
e5f4ede3f2 | ||
|
|
ce613bcf9d | ||
|
|
8ce8f0dfbe | ||
|
|
13c221244e | ||
|
|
a89b704fa1 | ||
|
|
0f228020ef | ||
|
|
8454d3d0a2 | ||
|
|
da15cc7531 | ||
|
|
5924aae1c0 | ||
|
|
46239012b4 | ||
|
|
f8119f3786 | ||
|
|
41c0598220 | ||
|
|
387382ec9f | ||
|
|
20998c8c31 | ||
|
|
f5d6edd9d0 | ||
|
|
4808a57289 | ||
|
|
96b907acbc | ||
|
|
d9c37ac763 | ||
|
|
146be2d0ad | ||
|
|
906375dc23 | ||
|
|
e1cdf03d40 | ||
|
|
8041da5424 | ||
|
|
7c4f881b57 | ||
|
|
81abc5adee | ||
|
|
331b7b6b13 | ||
|
|
fe23c59e0c | ||
|
|
666e1ebab2 | ||
|
|
3d05276312 | ||
|
|
00fdee5649 | ||
|
|
4a29e7bc99 | ||
|
|
a07019e3aa | ||
|
|
aa16fbaa8b | ||
|
|
9c15df8e22 | ||
|
|
111c70d6b4 | ||
|
|
0621f66889 | ||
|
|
4accd6b67e | ||
|
|
63ec110177 | ||
|
|
a27bb41b42 | ||
|
|
eaa78452bc | ||
|
|
4eb0a9c608 | ||
|
|
2e646f82b4 | ||
|
|
74bd40ba78 | ||
|
|
4f99efcf2c | ||
|
|
be10e78c7a | ||
|
|
b6a677a710 | ||
|
|
53d7784d5d | ||
|
|
10f2a0a8a4 | ||
|
|
eab5cc911c | ||
|
|
d88fe381db | ||
|
|
ecded4a8a7 | ||
|
|
772a3bdce2 | ||
|
|
fbf6a07da9 | ||
|
|
25e055a7ac | ||
|
|
e024a29157 | ||
|
|
83f3b7e226 | ||
|
|
d3d6f30f72 | ||
|
|
99c94ffaec | ||
|
|
1a836dac75 | ||
|
|
0447fa6f86 | ||
|
|
e8ee5b052b | ||
|
|
d7980da35e | ||
|
|
239e100b44 | ||
|
|
2216072669 | ||
|
|
8cf3ff31db | ||
|
|
b108a6fe62 | ||
|
|
f58ff68ab3 | ||
|
|
e7e9d9fecf | ||
|
|
d417bd5360 | ||
|
|
90f3b0b7da | ||
|
|
bce3bf61e9 | ||
|
|
49372996ac | ||
|
|
e3ed5a53f2 | ||
|
|
1231cd69a8 | ||
|
|
8e85a67e7b | ||
|
|
1badb93331 | ||
|
|
e23358ea0c | ||
|
|
73bbae14c1 | ||
|
|
74e3b8b067 | ||
|
|
9a1254048e | ||
|
|
7ae510ed28 | ||
|
|
620d7eb6f1 | ||
|
|
cd21fca749 | ||
|
|
ead5d38d7c | ||
|
|
c4bfa6bdfe | ||
|
|
d8a77c700b | ||
|
|
5fc53948c6 | ||
|
|
587440b6bc | ||
|
|
fb49eb20e8 | ||
|
|
60e87ba518 | ||
|
|
ff94ccb343 | ||
|
|
42711ff9ac | ||
|
|
7c957d96be | ||
|
|
427a31cdff | ||
|
|
51125fa05a | ||
|
|
24e347bb2d | ||
|
|
1a47f864f9 | ||
|
|
ad623353bc | ||
|
|
9d2610a911 | ||
|
|
bc9692f521 | ||
|
|
d8e3c09b57 | ||
|
|
7e7d845f61 | ||
|
|
b83c118c7f | ||
|
|
069815790b | ||
|
|
366e2a269f | ||
|
|
bdc2cdbadb | ||
|
|
fe58fb7826 | ||
|
|
6dbb076a1e | ||
|
|
3676526daf | ||
|
|
6a15a9d99a | ||
|
|
a46317de18 | ||
|
|
32355e0219 | ||
|
|
2cbc06a683 | ||
|
|
cce1e25c2b | ||
|
|
39dba3f8ab | ||
|
|
5bdf3bff60 | ||
|
|
65516ea3ac | ||
|
|
27989d066e | ||
|
|
3ae6298656 | ||
|
|
662f4f0d53 | ||
|
|
a02b918ff9 | ||
|
|
23825a824d | ||
|
|
cf67393db6 | ||
|
|
7ed0f5c23b | ||
|
|
f363c6cd55 | ||
|
|
c8a6f58007 | ||
|
|
93cdbaac57 | ||
|
|
47ad3b3075 | ||
|
|
1493cf2dcf | ||
|
|
35e59d2510 | ||
|
|
4f95cc6d13 | ||
|
|
f627658567 | ||
|
|
3c24550b19 | ||
|
|
98fab394a1 | ||
|
|
6fbc2b4125 | ||
|
|
64de77c065 | ||
|
|
19ba1e723c | ||
|
|
1d4dc44668 | ||
|
|
0fe3435033 | ||
|
|
6d9c856514 | ||
|
|
9634838f64 | ||
|
|
28273ffc2a | ||
|
|
7b65b64864 | ||
|
|
b90fd59544 | ||
|
|
d5ee2d7ef1 | ||
|
|
73358e6bae | ||
|
|
f2ad41e8e2 | ||
|
|
187dc188c2 | ||
|
|
ea2d3cf597 | ||
|
|
bd30c4e05a | ||
|
|
51283a626c | ||
|
|
01601d210b | ||
|
|
99b58c0c4f | ||
|
|
1001c25487 | ||
|
|
c0ace5a0ca | ||
|
|
1041c4bba8 | ||
|
|
3db594af86 | ||
|
|
71ba0a239f | ||
|
|
13de1dab82 | ||
|
|
8157673291 | ||
|
|
2697b46616 | ||
|
|
81aeb0619b | ||
|
|
693ac693a4 | ||
|
|
b22d034933 | ||
|
|
bb7c112a2d | ||
|
|
0101aef51a | ||
|
|
9eef5b50e6 | ||
|
|
2bda4cdaf8 | ||
|
|
b708a90548 | ||
|
|
90774f1442 | ||
|
|
a705aa36b7 | ||
|
|
60971f1af4 | ||
|
|
00dee37e22 | ||
|
|
2c16cc7116 | ||
|
|
5ad5c0df7e | ||
|
|
00328cbd0a | ||
|
|
7f22bb0612 | ||
|
|
7ebe8c0891 | ||
|
|
dbc0c8c6d2 | ||
|
|
04a0dcb39a | ||
|
|
24710af156 | ||
|
|
e149046272 | ||
|
|
1381a526ab | ||
|
|
fa7af04418 | ||
|
|
95050850d6 | ||
|
|
04aba81ab3 | ||
|
|
a11e3d73a7 | ||
|
|
793a75b8a2 | ||
|
|
976fdf6e50 | ||
|
|
a71723e51f | ||
|
|
7a3e8dd378 | ||
|
|
2cd1384786 | ||
|
|
9f32644c32 | ||
|
|
a7387c27fb | ||
|
|
2f8b7276a4 | ||
|
|
cbcfbfaaad | ||
|
|
a2debe972f | ||
|
|
57dad6d706 | ||
|
|
1f4b751fbb | ||
|
|
8808cbfc04 | ||
|
|
c6b56d580e | ||
|
|
b358dfcf58 | ||
|
|
57e2b0d670 | ||
|
|
b7df54e4f3 | ||
|
|
2ccc93bb14 | ||
|
|
2c1fd4c8bf | ||
|
|
466081aa66 | ||
|
|
be6bb3ffc5 | ||
|
|
17216a354d | ||
|
|
d1deac1bbb | ||
|
|
44275a8a04 | ||
|
|
e64d26af6c | ||
|
|
8ca52c64b9 | ||
|
|
ec3084db12 | ||
|
|
faa4a06c42 | ||
|
|
72d0e876c1 | ||
|
|
a04368e9fc | ||
|
|
a18f0cfc84 | ||
|
|
e1110b20df | ||
|
|
5b2ab93e60 | ||
|
|
533328aa91 | ||
|
|
6eb84c195a | ||
|
|
f8c4c93f20 | ||
|
|
e2b304daad | ||
|
|
c33f3c067b | ||
|
|
20c17648e6 | ||
|
|
240876e710 | ||
|
|
999dbcfdf3 | ||
|
|
67510b8fad | ||
|
|
24b526db93 | ||
|
|
dfe4c0d495 | ||
|
|
2b933c45dd | ||
|
|
f90b23ddd9 | ||
|
|
0eccd7b36e | ||
|
|
a95def9139 | ||
|
|
76ce424f8f | ||
|
|
6855dcb8fd | ||
|
|
e3ec54f213 | ||
|
|
aa3d92360e | ||
|
|
66f427dd2d | ||
|
|
25361fafdb | ||
|
|
081366a682 | ||
|
|
2100286dbb | ||
|
|
29e0daf80d | ||
|
|
57f3bf2807 | ||
|
|
775bdce6e1 | ||
|
|
b9b2a59915 | ||
|
|
33affc5a59 | ||
|
|
9adf760c8e | ||
|
|
eb24e60b8c | ||
|
|
9c744b34f8 | ||
|
|
6054f68754 | ||
|
|
74b579e281 | ||
|
|
c0da7d935d | ||
|
|
07f679b2e7 | ||
|
|
0c16c6db63 | ||
|
|
613634473a | ||
|
|
47861b0dc3 | ||
|
|
426ca2e5c7 | ||
|
|
91357a2d03 | ||
|
|
e4423200cc | ||
|
|
7d865b0a09 | ||
|
|
5b48561f36 | ||
|
|
a5be719261 | ||
|
|
77fb1abb85 | ||
|
|
b20a46ba81 | ||
|
|
9c150532fc | ||
|
|
4ab3387db4 | ||
|
|
92157aa5da | ||
|
|
801288c88c | ||
|
|
6f5c81c9a2 | ||
|
|
e5edffae4b | ||
|
|
3fac5751e4 | ||
|
|
f2d4d2fae1 | ||
|
|
e91a74d31f | ||
|
|
955f50b94e | ||
|
|
a8654280a7 | ||
|
|
083b5ec5ad | ||
|
|
500f9d92c8 | ||
|
|
67eabdc17c | ||
|
|
d9f917ecef | ||
|
|
eb4512aee3 | ||
|
|
0785a46fe3 | ||
|
|
78c9875d92 | ||
|
|
c135d46cbc | ||
|
|
4b6c5fe1cc | ||
|
|
66f9d6f435 | ||
|
|
4980b62dc2 | ||
|
|
d91fa89310 | ||
|
|
d6debc5408 | ||
|
|
c3ed7262cd | ||
|
|
455fb55857 | ||
|
|
1370a29233 | ||
|
|
bb5daad3cc | ||
|
|
6be765a9bf | ||
|
|
8eeb4c3d6c | ||
|
|
d0ff0c7d5c | ||
|
|
c431370a68 | ||
|
|
4f598f7422 | ||
|
|
7b1b017d8c | ||
|
|
1afeaf0e15 | ||
|
|
33e696071e | ||
|
|
44cf4e7516 | ||
|
|
2aa9302751 | ||
|
|
6cdb0d1eab | ||
|
|
d9225d06d9 | ||
|
|
37784bce6c | ||
|
|
073352b21e | ||
|
|
f4fd17be4d | ||
|
|
a011cd5e3b | ||
|
|
3a45c18555 | ||
|
|
38cd3edf99 | ||
|
|
e406bd445d | ||
|
|
f1603349cc | ||
|
|
d5c36db531 | ||
|
|
0f92609425 | ||
|
|
0f4d381ac0 | ||
|
|
ea0c713c89 | ||
|
|
4d00e37ca1 | ||
|
|
1e324636e0 | ||
|
|
df3c4bb369 | ||
|
|
2fefb56b8c | ||
|
|
934c0b56be | ||
|
|
e0cf83c944 | ||
|
|
9cabb04395 | ||
|
|
2ab7d3d98a | ||
|
|
228aae31c9 | ||
|
|
86a0ba55fe | ||
|
|
9e3755f9ca | ||
|
|
ab794db0ab | ||
|
|
b445ddc2c8 | ||
|
|
fb50f94ed8 | ||
|
|
b02857cf04 | ||
|
|
be306cf86c | ||
|
|
bf7a2cc911 | ||
|
|
4989486f45 | ||
|
|
9be6055113 | ||
|
|
8b0c2f6d1b | ||
|
|
141e2ca37c | ||
|
|
b803fefebd | ||
|
|
ba1979a662 | ||
|
|
1b46033584 | ||
|
|
7b2f79e3ac | ||
|
|
9145a66cef | ||
|
|
d5944c2443 | ||
|
|
d135dff5cd | ||
|
|
62a838ebdd | ||
|
|
f2d0e18f2c | ||
|
|
4347fc5f07 | ||
|
|
ce0cbd868c | ||
|
|
2658082607 | ||
|
|
55aa7f968f | ||
|
|
0882942164 | ||
|
|
917ff43795 | ||
|
|
d562644581 | ||
|
|
02aa73bb8d | ||
|
|
0fa871ef61 | ||
|
|
815ceef276 | ||
|
|
10c1705790 | ||
|
|
e467cdaf23 | ||
|
|
269ef9b2de | ||
|
|
657ee435d1 | ||
|
|
2f8af3a1b7 | ||
|
|
2c5e412d91 | ||
|
|
9d4a62fa5e | ||
|
|
714d97dd49 | ||
|
|
5a813eb4c8 | ||
|
|
185abdb442 | ||
|
|
60994c7667 | ||
|
|
eacea56d9d | ||
|
|
99486b1eaf | ||
|
|
aeb8c01370 | ||
|
|
2ad3a1bfed | ||
|
|
329bf25b14 | ||
|
|
90eee5674f | ||
|
|
1761087283 | ||
|
|
a13c5541cc | ||
|
|
262cd92300 | ||
|
|
87fa2693f0 | ||
|
|
64a28acd64 | ||
|
|
3ea80ef944 | ||
|
|
e1bbecad38 | ||
|
|
149d98a10b | ||
|
|
bd583b58dc | ||
|
|
7d0f7332b9 | ||
|
|
83508cfb1b | ||
|
|
969a85f303 | ||
|
|
961c57c6f0 | ||
|
|
89d6219e95 | ||
|
|
131c9ffc13 | ||
|
|
b1fdaced9f | ||
|
|
0cf208b80a | ||
|
|
9c8eb8a2bc | ||
|
|
452bf0e544 | ||
|
|
9d189e9dbf | ||
|
|
a3098ee67c | ||
|
|
73305a0cd1 | ||
|
|
e58f0352a7 | ||
|
|
8d393a1e1c | ||
|
|
d0b3588f6c | ||
|
|
7625e923e5 | ||
|
|
76ad744581 | ||
|
|
47110730e6 | ||
|
|
7f01b3040d | ||
|
|
f30f3263bd | ||
|
|
ca678b1784 | ||
|
|
ae7555a423 | ||
|
|
e0925b9d38 | ||
|
|
563ea0eb78 | ||
|
|
198532e280 | ||
|
|
69ffb5e1dc | ||
|
|
237156c46f | ||
|
|
264ba57cbb | ||
|
|
a537ae4217 | ||
|
|
8896e05e35 | ||
|
|
c617f64b6b | ||
|
|
eb36641523 | ||
|
|
885cceccae | ||
|
|
2d07c7d80e | ||
|
|
e02dc1be15 | ||
|
|
c34224effa | ||
|
|
987313e7dc | ||
|
|
6a959fb5e4 | ||
|
|
b00688bad1 | ||
|
|
ccc6c742ee | ||
|
|
0b4ff27be4 | ||
|
|
43b5443b30 | ||
|
|
76c4ca5ccf | ||
|
|
e2e76560a2 | ||
|
|
ce806ae854 | ||
|
|
c54039392d | ||
|
|
671bf2352b | ||
|
|
1ed6b92118 | ||
|
|
ab169a6f48 | ||
|
|
7a8fbbe06a | ||
|
|
53aa72d38c | ||
|
|
8a2707ee37 | ||
|
|
c377ddfcca | ||
|
|
5c4f8db497 | ||
|
|
132a657f00 | ||
|
|
7669cd75b8 | ||
|
|
3813884287 | ||
|
|
b80cf04cdc | ||
|
|
986353a0f0 | ||
|
|
e8b114094a | ||
|
|
cdce877601 | ||
|
|
77b24e9823 | ||
|
|
3f5ad22632 | ||
|
|
ce580d9935 | ||
|
|
3ae24939cc | ||
|
|
7162f16c1a | ||
|
|
19a9f6410d | ||
|
|
a1db8f7b2d | ||
|
|
d680bd0549 | ||
|
|
3084546b32 | ||
|
|
74ea1a5218 | ||
|
|
7505bffcee | ||
|
|
0061ddef72 | ||
|
|
19036621dc | ||
|
|
d96ddc8649 | ||
|
|
895347f20f | ||
|
|
5e92d9920a | ||
|
|
b2d6e1bcfd | ||
|
|
35f870c650 | ||
|
|
de12c1fada | ||
|
|
ec518d4758 | ||
|
|
9425d16c19 | ||
|
|
7252d4aad3 | ||
|
|
54919db13f | ||
|
|
2ce368abba | ||
|
|
2a09529e75 | ||
|
|
efa2f2edf8 | ||
|
|
004bde2c62 | ||
|
|
e528d23c68 | ||
|
|
43716a2c74 | ||
|
|
a7bf45faaf | ||
|
|
77845829ad | ||
|
|
ab35fcd84d | ||
|
|
d5d1403ebf | ||
|
|
7bf47be106 | ||
|
|
4f0cd1083a | ||
|
|
1e1129687f | ||
|
|
79059c20fa | ||
|
|
9c05692a24 | ||
|
|
e688500cf5 | ||
|
|
6f7bb8a6aa | ||
|
|
1777c56efa | ||
|
|
ca6f16bafb | ||
|
|
46ddc55345 | ||
|
|
10d6f372b3 | ||
|
|
de57f87b82 | ||
|
|
9dbcdc0bd7 | ||
|
|
6dc5c91f60 | ||
|
|
587fae1539 | ||
|
|
d2e071dece | ||
|
|
072df6b153 | ||
|
|
254b400caf | ||
|
|
976d8d84d7 | ||
|
|
bac1ef1c17 | ||
|
|
de2879bdee | ||
|
|
a018245f40 | ||
|
|
a22433967a | ||
|
|
9d27fac3b1 | ||
|
|
5866097d6a | ||
|
|
2c3f0b9cb1 | ||
|
|
fb13958881 | ||
|
|
c7ecf99c3f | ||
|
|
62bb166f9a | ||
|
|
59cee06f45 | ||
|
|
8bdf7eeb9c | ||
|
|
f30863f452 | ||
|
|
f75957130e | ||
|
|
ed4d194332 | ||
|
|
958391e326 | ||
|
|
2eb595552d | ||
|
|
c3562c2d0b | ||
|
|
1574db3eab | ||
|
|
ba07634a88 | ||
|
|
85c3aef1b0 | ||
|
|
35d7909828 | ||
|
|
0fb901e275 | ||
|
|
1d9907e0cf | ||
|
|
855bc53623 | ||
|
|
7fa0adbda4 | ||
|
|
084220692f | ||
|
|
abfbaa7f47 | ||
|
|
3d7b15d6bb | ||
|
|
223d5ebfc9 | ||
|
|
a6fe820a61 | ||
|
|
7e4db7b504 | ||
|
|
f7fa75fdfb | ||
|
|
0581c7b2f4 | ||
|
|
be6c16d74c | ||
|
|
c2597cb516 | ||
|
|
997f9d9117 | ||
|
|
6623d1e776 | ||
|
|
ef3980cf07 | ||
|
|
260f0a61ce | ||
|
|
69189bbf18 | ||
|
|
eb8c6bb1be | ||
|
|
2a062e5140 | ||
|
|
6fd22ae4ee | ||
|
|
93aa134aaa | ||
|
|
14ecf0c08a | ||
|
|
5d5b60a8ed | ||
|
|
5d32f58f5d | ||
|
|
b314aa4aff | ||
|
|
75f2a4e3fd | ||
|
|
cf180c1179 | ||
|
|
b4bc9267e9 | ||
|
|
99b2d1879b | ||
|
|
aa8698ddff | ||
|
|
76c0623190 | ||
|
|
eb64e52815 | ||
|
|
26fb409da5 | ||
|
|
2d76d10c9d | ||
|
|
0401d872e4 | ||
|
|
c9a0ad187b | ||
|
|
0aa5851549 | ||
|
|
ef7bc127b6 | ||
|
|
d8e3732f8d | ||
|
|
07b41559b6 | ||
|
|
6feb7ff3fa | ||
|
|
bafabe5fb3 | ||
|
|
ccaf19cfe4 | ||
|
|
97dfd522d0 | ||
|
|
f92ee4064b | ||
|
|
1003bee72a | ||
|
|
638aba20c0 | ||
|
|
fbe6a54e41 | ||
|
|
d058cf15c9 | ||
|
|
4de10a7d01 | ||
|
|
01e4f97361 | ||
|
|
284896fbd9 | ||
|
|
c4d4d8160d | ||
|
|
438de76655 | ||
|
|
86b8e894e5 | ||
|
|
719ca016f4 | ||
|
|
2b8571c52c | ||
|
|
585174ee94 | ||
|
|
3fb6207f53 | ||
|
|
bf3e7cc2c4 | ||
|
|
0564580605 | ||
|
|
7a12c3eee6 | ||
|
|
82810ab168 | ||
|
|
4749e5857c | ||
|
|
28d434b76c | ||
|
|
b27b2d62f5 | ||
|
|
7332eb81e0 | ||
|
|
ad98d41041 | ||
|
|
3c20c835da | ||
|
|
f42ecc8464 | ||
|
|
ac2986b141 | ||
|
|
e5169b241e | ||
|
|
12d04803bd | ||
|
|
f2a1a930dd | ||
|
|
7da9018378 | ||
|
|
978a5865a7 | ||
|
|
5d16b902cc | ||
|
|
554822fc6a | ||
|
|
11f3492d9d | ||
|
|
38916f8375 | ||
|
|
3c303210a3 | ||
|
|
5a7d5b5005 | ||
|
|
3c880b5bda | ||
|
|
6525b957c6 | ||
|
|
39498b71d0 | ||
|
|
e8608a97b7 | ||
|
|
373388a77b | ||
|
|
6153283de4 | ||
|
|
722e688783 | ||
|
|
39b8e47ff5 | ||
|
|
53d76f08c1 | ||
|
|
af76900aa6 | ||
|
|
3964f6fb46 | ||
|
|
4c37ab16fd | ||
|
|
08be751d23 | ||
|
|
df18657183 | ||
|
|
a2bcd58f78 | ||
|
|
55f8842dbd | ||
|
|
bfc3cbe204 | ||
|
|
a5292c4bb7 | ||
|
|
fdef61cc25 | ||
|
|
ba3eed39e3 | ||
|
|
420e13a746 | ||
|
|
181114aed5 | ||
|
|
b4d6866f40 | ||
|
|
ee5a07be1c | ||
|
|
3c900ae5c1 | ||
|
|
27bfc21cb0 | ||
|
|
fd1c4c7e3a | ||
|
|
ba1259c05c | ||
|
|
60ff658181 | ||
|
|
1bc76f8490 | ||
|
|
f036fbfe4a | ||
|
|
1f766fef31 | ||
|
|
f10f09eeb1 | ||
|
|
3937d5e33b | ||
|
|
129e8ae14a | ||
|
|
ca0d17ec2a | ||
|
|
30e7c2f6b0 | ||
|
|
819f5e6450 | ||
|
|
1e54752478 | ||
|
|
80b1ee4f0d | ||
|
|
af313352c0 | ||
|
|
aa92511ebc | ||
|
|
cac9d8e82f | ||
|
|
a24f9db2ff | ||
|
|
dab8d62e95 | ||
|
|
a535319f1d | ||
|
|
bd73e87bf9 | ||
|
|
273a2973e0 | ||
|
|
57773bcaf5 | ||
|
|
d05c5ac198 | ||
|
|
04cd15a22f | ||
|
|
4bb6bbf009 | ||
|
|
bac0f7c535 | ||
|
|
9c3ace95a7 | ||
|
|
17012bcd8c | ||
|
|
bdcdfea175 | ||
|
|
c441ac443a | ||
|
|
f9fe13bdc1 | ||
|
|
415c260175 | ||
|
|
437c525fbd | ||
|
|
a5936331c9 | ||
|
|
d9ed8c7704 | ||
|
|
0ded0c87a1 | ||
|
|
d70c682be2 | ||
|
|
0b3d7d542f | ||
|
|
c1919febc2 | ||
|
|
127b070f7b | ||
|
|
369a9e6c19 | ||
|
|
de50a01ab2 | ||
|
|
b5d33cd6a5 | ||
|
|
0bcb773332 | ||
|
|
ce0b3fb74b | ||
|
|
6aaf04e25b | ||
|
|
b7894681e6 | ||
|
|
a41adfe020 | ||
|
|
7178be73a5 | ||
|
|
a55a3afad5 | ||
|
|
2223cf80f7 | ||
|
|
6c0164713b | ||
|
|
13ed584d73 | ||
|
|
6104cde755 | ||
|
|
7d8499686a | ||
|
|
3cdb65cf8e | ||
|
|
0cb429222a | ||
|
|
e9ea67ec14 | ||
|
|
cb1c5d33cb | ||
|
|
abb4968487 | ||
|
|
65075ed599 | ||
|
|
9655a30d25 | ||
|
|
de6fcc3997 | ||
|
|
8be54ab5ee | ||
|
|
1a8d40d631 | ||
|
|
3e71c314d7 | ||
|
|
08fc551d9f | ||
|
|
bc70ebe5ee | ||
|
|
8f696d85f7 | ||
|
|
e18ee9fcf0 | ||
|
|
b316e1b49f | ||
|
|
74dff404b8 | ||
|
|
d8c456a8c9 | ||
|
|
9f82eb5678 | ||
|
|
5e89f334bb | ||
|
|
f8325a7067 | ||
|
|
cb7d7a688d | ||
|
|
65f6bd7b8a | ||
|
|
60535ba842 | ||
|
|
e0ae80d2ea | ||
|
|
98d2e6efa4 | ||
|
|
4effb55756 | ||
|
|
d22b18bf9f | ||
|
|
bc81e5a9e0 | ||
|
|
492f268b10 | ||
|
|
68323f9c4a | ||
|
|
4b28529736 | ||
|
|
d3454a3dab | ||
|
|
ee26eebc6b | ||
|
|
ec2073346e | ||
|
|
4ade80d9cb | ||
|
|
63b1f0a76c | ||
|
|
e9b9d1b5da | ||
|
|
a8c5b2af29 | ||
|
|
212b61ce15 | ||
|
|
3e9c7b4921 | ||
|
|
3e89997ab5 | ||
|
|
67e274fd4e | ||
|
|
55e817d234 | ||
|
|
3b645b2108 | ||
|
|
279e6be282 | ||
|
|
52cc6b4721 | ||
|
|
f0f136f28a | ||
|
|
0319b13304 | ||
|
|
89df1a98db | ||
|
|
bea7b6b045 | ||
|
|
586c7c36c1 | ||
|
|
beac98427d | ||
|
|
1ea3f595cc | ||
|
|
4722382d49 | ||
|
|
a250d590ca | ||
|
|
753f99ac0c | ||
|
|
a943a5c294 | ||
|
|
b093cae70d | ||
|
|
4187d52bac | ||
|
|
46b27f609a | ||
|
|
fdba23c32d | ||
|
|
55a38b6b74 | ||
|
|
44f33da454 | ||
|
|
070f9caf7f | ||
|
|
9d6772fbd6 | ||
|
|
cce6fc5b65 | ||
|
|
f6fb3add77 | ||
|
|
9c68429e25 | ||
|
|
b9b1d20499 | ||
|
|
cbf8593225 | ||
|
|
cd2fd1af70 | ||
|
|
e0fdaa1868 | ||
|
|
765442c25d | ||
|
|
da699b54fa | ||
|
|
5234b79210 | ||
|
|
ad1890109a | ||
|
|
fb0530deba | ||
|
|
9191a54637 | ||
|
|
2989cefb79 | ||
|
|
1956b6b068 | ||
|
|
5bd8d13e66 | ||
|
|
e449e1f954 | ||
|
|
4449bfc646 | ||
|
|
95fd38799d | ||
|
|
4aaa3c507a | ||
|
|
192664605b | ||
|
|
c97a30b9c4 | ||
|
|
7d4994f317 | ||
|
|
d1ee880586 | ||
|
|
b59b94099a | ||
|
|
dfc24e37c0 | ||
|
|
c7e81edf4d | ||
|
|
be5332d0a4 | ||
|
|
a5755921ab | ||
|
|
3dd867c684 | ||
|
|
1e9d8bc75b | ||
|
|
084dc819d9 | ||
|
|
d56b02e410 | ||
|
|
da3d402f1a | ||
|
|
9a26478498 | ||
|
|
5edb32f48d | ||
|
|
7e01f6d3db | ||
|
|
10fc703e57 | ||
|
|
d09ea11694 | ||
|
|
0121462576 | ||
|
|
974e3cf2cf | ||
|
|
64f2c89a40 | ||
|
|
af89258bbd | ||
|
|
1e9681a71e | ||
|
|
6549ed0301 | ||
|
|
5ebf1fc96f | ||
|
|
b5d7413d88 | ||
|
|
da0c303f22 | ||
|
|
26f3778786 | ||
|
|
e6a0ab8e64 | ||
|
|
0fe99cf2a9 | ||
|
|
2d5a084d2b | ||
|
|
afe5c10679 | ||
|
|
921900b23b | ||
|
|
788c04ce9d | ||
|
|
1184be3e28 | ||
|
|
23062acdac | ||
|
|
26693fd11f | ||
|
|
ce9748096c | ||
|
|
d2eb060fe5 | ||
|
|
86048b32d3 | ||
|
|
8f05f0d27c | ||
|
|
00a8a64c4e | ||
|
|
e95fd30e13 | ||
|
|
fb051b60c1 | ||
|
|
64d95aa991 | ||
|
|
446794e41d | ||
|
|
072c8b720d | ||
|
|
6078af0dbb |
+57
@@ -0,0 +1,57 @@
|
||||
# ── Personal configuration — never push to public repo ────────────────────────
|
||||
# These contain API keys, passwords, SSH paths, personal hostnames.
|
||||
# Templates (*.template) are safe and remain tracked.
|
||||
Configurations/host*.conf
|
||||
Configurations/master.conf
|
||||
# *.bak alone missed conf_upgrade's real output — it writes host1.conf.bak-20260802, which does
|
||||
# not end in .bak — so those sat untracked rather than ignored, one `git add -A` from being
|
||||
# pushed. The glob has to cover the suffix.
|
||||
Configurations/*.bak*
|
||||
.vscode
|
||||
|
||||
# ── Personal scratch notes — dev-only, never pushed ───────────────────────────
|
||||
Notes_To-Do.md
|
||||
|
||||
# ── Runtime state, data, logs ─────────────────────────────────────────────────
|
||||
# Contents, not the directory itself. Ignoring "data/" outright means git never descends into
|
||||
# it, and a negation for a file inside an excluded directory is silently ineffective — so the
|
||||
# README explaining what data/ is would be the one file missing from every installation of it.
|
||||
data/*
|
||||
!data/README.md
|
||||
State_Files/
|
||||
.cache/
|
||||
*.log
|
||||
*.lock
|
||||
|
||||
# ── Plugin runtime state (per-host, never synced) ────────────────────────────
|
||||
schedule.json
|
||||
docker_folders.json
|
||||
*.db
|
||||
|
||||
# ── Plugin config files (co-located with scripts on flash) ───────────────────
|
||||
varaverk.cfg
|
||||
varaverk.cron
|
||||
varaverk-*.txz
|
||||
|
||||
# ── Build artifacts ───────────────────────────────────────────────────────────
|
||||
# .txz packages are attached to GitHub releases, not committed to the repo.
|
||||
Plugin/dist/
|
||||
|
||||
|
||||
# ── OS / editor ───────────────────────────────────────────────────────────────
|
||||
.DS_Store
|
||||
*.swp
|
||||
*~
|
||||
.vscode/
|
||||
|
||||
# ── Local-only plugin surfaces (per-installation, never pushed) ───────────────
|
||||
# Varaverk.page discovers pages/local/*.php and registers each as a tab; api/local/ holds their
|
||||
# endpoints. Both are symlinks into a store outside this repo, so what they contain belongs to
|
||||
# one installation and is not part of the project. The tracked loader is deliberately generic —
|
||||
# it names no page — so the public mirror never learns what a given server runs here.
|
||||
#
|
||||
# No trailing slash on either pattern. These paths are symlinks, not directories, and git treats
|
||||
# a symlink as a blob — a "dir/" pattern does not match one, so the entries sat untracked rather
|
||||
# than ignored, which is the same near-miss the *.bak rule above documents.
|
||||
Plugin/unraid/pages/local
|
||||
Plugin/unraid/api/local
|
||||
Vendored
-10
@@ -1,10 +0,0 @@
|
||||
{
|
||||
"files.exclude": {
|
||||
"**/.cache/**": true,
|
||||
"**/.next/**": true,
|
||||
"**/build/**": true,
|
||||
"**/coverage/**": true,
|
||||
"**/dist/**": true,
|
||||
"**/node_modules/**": true
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
# The index is generated, host-specific, and regenerable in minutes. Never commit it.
|
||||
*.db
|
||||
*.db-wal
|
||||
*.db-shm
|
||||
+370
@@ -0,0 +1,370 @@
|
||||
# ━━━━━ AI ━━━━━
|
||||
|
||||
Retrieval over Varaverk's own documentation. Ask the system a question about itself and get an
|
||||
answer grounded in its actual headers, READMEs, Manuals and conf templates — with sources.
|
||||
|
||||
Everything here is **off by default and optional**. Varaverk works exactly as well with
|
||||
`AI_ENABLED=false` as with it true. Nothing in the ecosystem depends on this folder.
|
||||
|
||||
| File | Role |
|
||||
|------|------|
|
||||
| `ai_index.sh` | Build / refresh the retrieval index |
|
||||
| `ai_query.sh` | Ask a question, or search the index directly |
|
||||
| `lib/chunk.js` | Split repo files into retrieval units |
|
||||
| `lib/index.js` | Embed chunks, store vectors in SQLite |
|
||||
| `lib/search.js` | Embed a query, score it, rank results |
|
||||
| `lib/cli.js` | Argument bridge the bash entry points call |
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHY THIS WORKS AT ALL ━━━
|
||||
|
||||
**Because the header audit came first.**
|
||||
|
||||
The single worst failure in naive retrieval is a chunk that contains half of one idea and half
|
||||
of another — a fixed-size window cutting mid-thought, embedding two unrelated things as one
|
||||
vector. That problem does not exist here, because every script in the repo carries the same six
|
||||
sections at exact, greppable boundaries:
|
||||
|
||||
```
|
||||
PURPOSE → OPERATIONAL MODEL → DESIGN PRINCIPLES → OPERATIONAL SAFEGUARDS
|
||||
→ CONFIGURATION → RUNTIME MODES
|
||||
```
|
||||
|
||||
Split on those and every chunk is a coherent unit by construction. No token windows, no overlap
|
||||
heuristics, no tuning.
|
||||
|
||||
**And because the headers say *why*.** A model can read `mover_stop.sh` and describe what it
|
||||
does. It cannot look at that code and know the cache writers are lockless *on purpose*, or that
|
||||
`removeCompletedDownloads` being true on both arrs is intentional. Those live in
|
||||
`DESIGN PRINCIPLES` and `OPERATIONAL SAFEGUARDS`, which is precisely what makes this index worth
|
||||
more than an equivalent pile of source.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SECTION ROUTING ━━━
|
||||
|
||||
Every chunk stores its section name as its own column, so a question's *shape* can steer
|
||||
retrieval before similarity is even considered:
|
||||
|
||||
| Question shape | Steered toward |
|
||||
|----------------|----------------|
|
||||
| "what stops X and Y overlapping" | `OPERATIONAL SAFEGUARDS` |
|
||||
| "which variable controls X" | `CONFIGURATION` |
|
||||
| "does this take --dry-run" | `RUNTIME MODES` |
|
||||
| "why is it built this way" | `DESIGN PRINCIPLES` |
|
||||
| "how does X work" | `OPERATIONAL MODEL` |
|
||||
|
||||
Applied as a **score boost, not a filter**. Intent detection is a heuristic, and a heuristic
|
||||
must never be able to exclude the one chunk that holds the answer. `--section=NAME` forces a
|
||||
hard filter when you actually want one.
|
||||
|
||||
The boost still has a blind spot worth knowing about: **definitional questions.** "What is
|
||||
Varaverk?" matches the `PURPOSE` intent, so every script's one-line PURPOSE gets boosted above
|
||||
the top-level prose that actually answers it — and the model correctly replies that the context
|
||||
does not define the system. The corpus is fine; `README.md` is indexed. The routing simply
|
||||
buries it. `--kind=readme` is the hard filter for that case:
|
||||
|
||||
```bash
|
||||
bash AI/ai_query.sh --kind=readme "what is Varaverk"
|
||||
```
|
||||
|
||||
`--kind` filters on where a chunk came from — `header`, `readme`, `manual`, `template`, `doc`, `ui` —
|
||||
and composes with `--section`. Prefer it over `--section` for "what is" and "why does this
|
||||
exist" questions, where the answer is narrative rather than a header field.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ NAMED-PARAGRAPH SUB-CHUNKING ━━━
|
||||
|
||||
Sections alone were not granular enough, and the failure was instructive.
|
||||
|
||||
`rsync.sh` documents fourteen distinct safeguards in one 2.8k-character
|
||||
`OPERATIONAL SAFEGUARDS` block. Asked *"what happens if pass 1 of a merge run fails"*, the
|
||||
correct chunk scored **0.558** — below unrelated chunks from other files — because the other
|
||||
thirteen safeguards dominated the vector.
|
||||
|
||||
The header convention writes each safeguard as a named paragraph: an unindented title, an
|
||||
indented body. Splitting on those titles took the same query to **0.718** and first place.
|
||||
|
||||
```
|
||||
Merge-Run Delete Interlock ← its own chunk
|
||||
--delete is applied only when pass 1 completed...
|
||||
```
|
||||
|
||||
The parent section name is carried onto every sub-chunk, so routing still works. The title
|
||||
detection requires the *next* line to be indented — without that check, any wrapped prose line
|
||||
became a spurious boundary mid-sentence.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHAT IS INDEXED ━━━
|
||||
|
||||
Roughly 2,900 chunks across ~180 files:
|
||||
|
||||
| Kind | Source |
|
||||
|------|--------|
|
||||
| `header` | Script headers — bash and PHP, sections and named paragraphs |
|
||||
| `readme` | Every `README-*.md` |
|
||||
| `manual` | Every `Manual-*.md` |
|
||||
| `template` | `Deployment/*.template` — the versioned conf schema |
|
||||
| `doc` | Top-level `README.md`, `Manual.md`, design notes |
|
||||
| `ui` | `Plugin/unraid/pages/readme/*.md` — the WebGUI's own help panels |
|
||||
|
||||
**Script bodies are not indexed.** Headers state intent, code states mechanism; for the
|
||||
questions this answers, intent retrieves better and costs far less.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE SAFETY BOUNDARY ━━━
|
||||
|
||||
**Only `git ls-files` is ever indexed.** This is not a convenience — it is the security model.
|
||||
|
||||
`Configurations/` and `data/` are gitignored, so every file holding a credential
|
||||
was never in the repo to begin with. The index therefore describes the full conf schema (via the
|
||||
tracked templates, which carry all the explanatory comments) while structurally **never
|
||||
containing a secret**.
|
||||
|
||||
> Do not "improve" this into a filesystem walk. A logged secret can be rotated. A secret
|
||||
> averaged into a 768-dimension float cannot be found, let alone removed.
|
||||
|
||||
A live conf value the model genuinely needs should arrive through a tool call at query time,
|
||||
subject to redaction — never baked into a vector at index time.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ HOW IT IS BUILT ━━━
|
||||
|
||||
**SQLite, no vector database.** ~2,900 chunks × 768 dims is a couple of million multiply-adds
|
||||
per query — under a millisecond. A vector DB would be a container to run, monitor, fail over and
|
||||
back up, in exchange for nothing at this scale.
|
||||
|
||||
**Vectors are raw little-endian float32 BLOBs.** `nomic-embed-text` returns L2-normalised
|
||||
vectors, so cosine similarity is a plain dot product — no normalising, no magnitude cache. PHP
|
||||
reads the same blobs with `unpack('f*', $blob)` when the UI needs them.
|
||||
|
||||
**Incremental on mtime.** A file whose mtime has not moved is skipped without being read. A
|
||||
no-op refresh takes about 70 ms; a full rebuild takes a few minutes.
|
||||
|
||||
**Node for the maths, bash for everything else.** The bash entry points own configuration,
|
||||
gating, locking and logging exactly as every other Varaverk job does. Node owns only float
|
||||
vector maths and SQLite BLOBs. Same split as `api_cache_writer.sh` and its PHP.
|
||||
|
||||
> `lib/index.js` uses `node:sqlite`, which Node still marks experimental. It is used because it
|
||||
> needs no native compilation on Unraid. If a Node upgrade ever breaks it, the index is
|
||||
> regenerable in minutes — this is a disposable artefact, not a datastore.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ USAGE ━━━
|
||||
|
||||
```bash
|
||||
# Build (needs AI_ENABLED=true)
|
||||
bash AI/ai_index.sh # incremental
|
||||
bash AI/ai_index.sh --force # full rebuild
|
||||
bash AI/ai_index.sh --dry-run # what would be indexed; contacts nothing
|
||||
bash AI/ai_index.sh --status # size, counts, last build
|
||||
|
||||
# Ask
|
||||
bash AI/ai_query.sh "what stops rsync and the mover running at once"
|
||||
bash AI/ai_query.sh --search "why are the cache writers lockless"
|
||||
bash AI/ai_query.sh --section=CONFIGURATION "which variable sets the mover grace period"
|
||||
bash AI/ai_query.sh --kind=readme "what is Varaverk" # definitional / narrative
|
||||
bash AI/ai_query.sh --json "..." # for other scripts
|
||||
```
|
||||
|
||||
**`--search` is the trustworthy mode.** It returns verbatim repo text with nothing generated.
|
||||
When the answer matters, use it — or read the sources the generated answer cites.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ TRUST THE SOURCES, NOT THE PROSE ━━━
|
||||
|
||||
The generation prompt instructs the model to answer only from retrieved context and to say what
|
||||
is missing rather than fill the gap from general knowledge. That instruction matters here more
|
||||
than usual: this repo's conventions are frequently *not* the conventional ones, and a confident
|
||||
generic answer about rsync, Docker or systemd is worse than no answer.
|
||||
|
||||
It works — and its faithfulness cut both ways on the first real test. Asked which variable
|
||||
controls the mover's grace window, the model answered `MOVER_STOP_TIMEOUT`, "defaults to 30
|
||||
seconds", citing `mover_stop.sh › CONFIGURATION`. The variable was right. The 30 was wrong — the
|
||||
real value is 300 — and the model was quoting the header verbatim. **The header was stale.**
|
||||
|
||||
That sweep then found six stale `(default: N)` claims across the repo, all since corrected. The
|
||||
lesson is the operating principle for this whole folder:
|
||||
|
||||
> Retrieval is exactly as accurate as the documentation it points at. When an answer looks
|
||||
> wrong, check the cited source before blaming the model — it is usually reporting a real
|
||||
> problem in the repo.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHAT THIS DOES NOT DO ━━━
|
||||
|
||||
- **It does not write conf.** `AI_CONF_WRITE_ENABLED` exists in `master.conf` and is off, with
|
||||
an empty key whitelist. Nothing in this folder writes a setting.
|
||||
- **It does not make decisions.** No watchdog, cleanup or fallback path consults it. The
|
||||
per-feature `AI_ASSIST_*` toggles are all false and each one is earned separately.
|
||||
- **It does not index code.** Headers and docs only.
|
||||
- **It does not sync.** The index is host-local and gitignored. Each node builds its own.
|
||||
- **It is not required.** Every script runs identically with `AI_ENABLED=false`.
|
||||
|
||||
See `Notes_AI-Design.md` at the repo root for the wider design — host resolution across the
|
||||
Tailscale mesh, per-feature rollout tiers, and the conf-write guardrails that would have to be
|
||||
built before any of that is enabled.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SCHEDULING ━━━
|
||||
|
||||
The index tracks the repo automatically, from the only moment the corpus actually changes — a
|
||||
successful pull. `git_pull_execute.sh` re-indexes behind three gates: the pull succeeded and
|
||||
changed tracked files, `AI_INDEX_ON_PULL=true`, and `AI_ENABLED=true`. It is never fatal — a
|
||||
failed index leaves the previous one in place and the pull still reports success.
|
||||
|
||||
No cron entry and no `DAILY_MAINTENANCE_SCRIPTS` line are needed; the daily pull carries it.
|
||||
|
||||
An incremental run on an unchanged repo is ~70 ms, so a daily entry costs effectively nothing
|
||||
and a pull that changed twelve files costs a few seconds.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ PROFILES ━━━
|
||||
|
||||
A profile is a contract plus a set of inputs. `Plugin/unraid/include/ai_profiles.php` is the one
|
||||
definition of both, read by the endpoint, the worker, the shared chat include and the Scheduler
|
||||
dock.
|
||||
|
||||
| Profile | Turns | Retrieves | Notes |
|
||||
|---|---|---|---|
|
||||
| `varaverk` | 3 | yes | answers only from the index, with citations. The default. |
|
||||
| `chat` | 8 | **no** | ordinary conversation. Holds zero capabilities, deliberately. |
|
||||
| `code` | 4 | no | drafts shell for Custom Scripts; scans its own output for destructive ops |
|
||||
| `troubleshoot` | 2 | yes | reasons from evidence — an open log, or one you name. May file bug reports. |
|
||||
|
||||
Capabilities are granted per profile — retrieval, live health, run evidence, scoped log,
|
||||
incidents, conf lookup, bug filing, code scanning. `chat` holding an empty list is a guarantee,
|
||||
not an oversight: anything added to it stops being general chat and becomes an assistant that
|
||||
sometimes invents claims about this installation.
|
||||
|
||||
### Routing out of General Chat
|
||||
|
||||
`chat` hands a question to whichever profile fits, decided by `vv_ai_route_from_chat()`. Ordered
|
||||
most specific first, because these overlap on purpose:
|
||||
|
||||
| Question | Goes to | Why |
|
||||
|---|---|---|
|
||||
| "write me a script that prunes logs" | `code` | asked for something written |
|
||||
| "why did the daily orch fail" | `troubleshoot` | diagnostic phrasing **and** something here to diagnose |
|
||||
| "how did the daily orch go" | `varaverk` | about this install, but not a fault |
|
||||
| "what does arr_sync.sh do" | `varaverk` | names a script, wants documentation |
|
||||
| "why is the sky blue" | stays `chat` | diagnostic phrasing about nothing here |
|
||||
|
||||
`code` is checked first because it is the only intent about a thing that does not exist yet, so
|
||||
nothing else can claim it — and it is anchored on the verb, which is what keeps "write me a
|
||||
script" apart from "what does this script do". Escalation adds capability, so a wrong escalation
|
||||
costs more than a missed one: anything unrecognised stays in `chat`, the profile that cannot
|
||||
invent claims about this system. The worker reverts to `chat` anyway if retrieval comes back
|
||||
empty.
|
||||
|
||||
The answer opens with one line naming the profile that took it, because the button still shows
|
||||
the one you picked and an answer arriving under a different contract otherwise reads as the
|
||||
assistant ignoring you.
|
||||
|
||||
Routing is asserted by `Plugin/unraid/Tools/ai_explain_check.sh` against
|
||||
`ai_explain_fixtures.txt` — every case runs through the worker's `--explain` mode, which stops
|
||||
where deterministic assembly ends and never calls the model.
|
||||
|
||||
This used to live in five places — history depth in the endpoint, capabilities in `include/ai.php`,
|
||||
label and depth again in JavaScript, a prompt branch in the worker, and a label map on the
|
||||
Scheduler page. They had already drifted: the JavaScript knew three profiles where PHP knew four.
|
||||
The system prompts still live in `Tools/ai_chat_worker.php`, because they have one reader and
|
||||
moving them would relocate the most delicate text in the subsystem without removing a duplicate.
|
||||
|
||||
## ━━━ CONVERSATIONS ━━━
|
||||
|
||||
Chats are stored server-side under `AI_DATA_DIR/ai_chats/`, one JSON file each, saved
|
||||
automatically when a turn completes and pruned to `AI_CHAT_HISTORY_MAX` (default 10, oldest
|
||||
first by creation).
|
||||
|
||||
There is no Save button. A conversation worth keeping is not reliably one you knew was worth
|
||||
keeping while you were having it.
|
||||
|
||||
The same store backs the AI tab and the Monitor tab's AI row, so a thread started on the
|
||||
dashboard is the one you carry on in the tab. Messages are re-validated per message on the way
|
||||
in — a stored chat is replayed into a later prompt when reopened, so an unchecked role written
|
||||
there would be an injection that survives a reload rather than one turn.
|
||||
|
||||
Reopened chats render as plain turns: sources, reasoning and timings describe one generation and
|
||||
are not stored, because redrawing them beside a transcript that may be continued under a
|
||||
different profile would be citing evidence for an answer no longer being made.
|
||||
|
||||
### Secrets are redacted on the way to disk
|
||||
|
||||
A conversation about settings is a conversation containing credentials — asking for an API key to
|
||||
be changed means typing one. Message bodies are redacted in `vv_ai_chat_save()`, and the question
|
||||
is redacted again before it reaches `ai.log`.
|
||||
|
||||
**On the way out, never in flight.** The live turn keeps the real value, because the model needs
|
||||
it to carry out what was asked. What it does not need is that value still in the transcript a week
|
||||
later — and a stored chat is replayed into a later prompt when reopened, so an unredacted one
|
||||
would hand the credential back on every subsequent turn, indefinitely.
|
||||
|
||||
Two passes, because they catch different things:
|
||||
|
||||
| Pass | Catches | Method |
|
||||
|---|---|---|
|
||||
| Known values | a credential this host already holds | exact match against secret-shaped conf keys, longest first |
|
||||
| Assignment shapes | a credential arriving that is not in the conf yet | `NAME=value`, `"api_key": value`, "set the token to …" |
|
||||
|
||||
The second pass is the one that matters for settings changes: *"change the Emby API key to X"* is
|
||||
a secret arriving, and X matches nothing on disk until after the write it is requesting.
|
||||
|
||||
Ordinary prose is left alone — the patterns anchor on a secret-shaped *name*, so `CACHE_WARN_GB=100`
|
||||
and "turn off the zfs scrub" pass through untouched. `vv_conf_key_is_secret()` is shared with the
|
||||
conf audit log, so the two cannot disagree about what counts as a secret.
|
||||
|
||||
## ━━━ TOKEN ACCOUNTING ━━━
|
||||
|
||||
Every completed `ask` appends one row to `AI_TOKEN_DB` (`data/ai/ai_token_history.db`):
|
||||
|
||||
```
|
||||
date|time|host|profile|source|prompt_tokens|completion_tokens|tok_s
|
||||
2026-08-04|22:03:51|host1|varaverk|cli|2041|318|61.4
|
||||
```
|
||||
|
||||
Both paths write it — this CLI (`source=cli`) and the WebGUI worker (`source=webgui`) — so the
|
||||
totals are not quietly the tab's alone. `ai_query.sh` passes `--token-db` and `--token-host`;
|
||||
called by hand without them, `cli.js` simply skips the row rather than guessing a path, because
|
||||
this file never reads conf itself.
|
||||
|
||||
Read it on the plugin's AI tab, which aggregates today / last 7 days / all time, per host. Or
|
||||
straight from the shell, since it is just a delimited file:
|
||||
|
||||
```bash
|
||||
# tokens used today
|
||||
awk -F'|' -v d="$(date +%F)" '$1==d {p+=$6; c+=$7} END {print p+c}' data/ai/ai_token_history.db
|
||||
```
|
||||
|
||||
**The host column is where the turn ran, not where the file is read.** Each host writes only its
|
||||
own rows.
|
||||
|
||||
`ai_token_sync.sh` pulls each partner's ledger into `$AI_TOKEN_CACHE_DIR/<slot>.tokens.db` (`/tmp/varaverk/ai/`) — the
|
||||
same trick `conf_sync.sh` uses for partner confs, and it runs from
|
||||
`INTERMEDIATE_MAINTENANCE_SCRIPTS` every four hours. The tab then reads every ledger it can see,
|
||||
so a fleet total is a fleet total.
|
||||
|
||||
Pull only, no push: nothing here is needed by anyone else, and a reader that fetches its own data
|
||||
controls its own freshness instead of depending on the partner's cron. A partner file may only
|
||||
contribute rows whose `host` column matches its filename — a ledger copied into the wrong slot
|
||||
would otherwise be double-counted against a total that still looked plausible.
|
||||
|
||||
The cache is tmpfs with **no save/restore pair**, unlike the conf cache. Stale counters are worse
|
||||
than absent ones: absent renders as "not collected here", stale renders as fact. An unreachable
|
||||
partner leaves its file alone and logs at info, because a partner being down for weeks is a
|
||||
normal state, not an incident.
|
||||
|
||||
Pruning is by row count (`AI_TOKEN_RETAIN_ROWS`, default 20000) and happens on write, but only
|
||||
once the file passes a size threshold — an ordinary turn costs a `stat()` and an append. The CLI
|
||||
deliberately does not prune: duplicating a read-modify-write of the whole file in a second
|
||||
language is how the two drift apart.
|
||||
Executable
+265
@@ -0,0 +1,265 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================== AI Retrieval Index Builder ========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ==============================================================================================
|
||||
# Builds and refreshes the retrieval index over Varaverk's own documentation — every script
|
||||
# header, folder README and Manual, and the conf templates. Chunks each file on the section
|
||||
# boundaries the header convention already defines, embeds each chunk through Ollama, and
|
||||
# stores the vectors in SQLite for AI/ai_query.sh to search.
|
||||
#
|
||||
# The index is derived data. It is gitignored, host-local, and rebuildable from the repo in
|
||||
# minutes — nothing depends on it surviving.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
# Chunking mirrors the header convention rather than using a fixed token window:
|
||||
#
|
||||
# bash headers split on the six section names, then sub-split named-paragraph
|
||||
# safeguards and principles so one question finds one answer
|
||||
# PHP headers same six names plus the per-layer tails (EXPORTS, RENDERS, ...)
|
||||
# markdown split on headings
|
||||
# conf templates split on the ━━━ section rules
|
||||
#
|
||||
# Every chunk keeps its section name as a field, which is what lets a query about a safeguard
|
||||
# be steered toward OPERATIONAL SAFEGUARDS chunks before similarity is considered.
|
||||
#
|
||||
# Incremental by file mtime. A file whose mtime has not moved is skipped without being read,
|
||||
# so a routine refresh costs well under a second and a full rebuild costs a few minutes.
|
||||
#
|
||||
# Heavy lifting runs in Node — float vector maths and SQLite BLOBs are genuinely awkward in
|
||||
# bash. This follows the api_cache_writer.sh precedent: a bash shim owning config, gating,
|
||||
# locking and logging, in front of the language that fits the work.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Tracked Files Only
|
||||
# Indexes exactly what `git ls-files` reports. Configurations/, State_Files/ and data/ are
|
||||
# gitignored, so it is structurally impossible for a credential to reach the index — the
|
||||
# files holding them were never in the repo. This must never become a filesystem walk: a
|
||||
# secret written into a vector cannot be rotated back out of it.
|
||||
#
|
||||
# The Index Is Disposable
|
||||
# Stored under DATA_DIR, gitignored, and never synced to a partner. Losing it costs one
|
||||
# rebuild. Nothing reads it as a source of truth — it points at files, and the files are
|
||||
# the truth.
|
||||
#
|
||||
# Documentation Is The Corpus, Not The Code
|
||||
# Script bodies are not indexed. The headers state intent and the code states mechanism;
|
||||
# for the questions this answers, intent retrieves far better and is far cheaper.
|
||||
#
|
||||
# Off By Default
|
||||
# Does nothing unless AI_ENABLED is true. A node with AI off never pays for this.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Fail-Closed Gate
|
||||
# Exits cleanly unless AI_ENABLED is exactly "true". Any other value, including unset,
|
||||
# means off.
|
||||
#
|
||||
# Root Enforcement
|
||||
# Writes into DATA_DIR alongside other Varaverk state.
|
||||
#
|
||||
# Concurrency Lock
|
||||
# acquire_lock() prevents two indexers racing on the same database.
|
||||
#
|
||||
# Dependency Verification
|
||||
# Verifies node and the Ollama endpoint before touching the database. A missing dependency
|
||||
# is reported and exits non-zero rather than leaving a half-built index.
|
||||
#
|
||||
# Reachability Pre-flight
|
||||
# Probes the resolved Ollama URL with AI_CONNECT_TIMEOUT before starting. An unreachable
|
||||
# endpoint aborts immediately instead of failing once per batch across the whole corpus.
|
||||
#
|
||||
# Partial Failure Is Not Recorded As Success
|
||||
# A file whose chunks all failed to embed keeps its previous rows and its old mtime, so the
|
||||
# next run retries it. A run with any failed batch exits 3.
|
||||
#
|
||||
# Atomic Per-Run Write
|
||||
# All database changes commit in one transaction. An interrupted run leaves the previous
|
||||
# index intact rather than a partially rewritten one.
|
||||
#
|
||||
# Deleted Files Are Removed From The Index
|
||||
# A file that has left the repo has its chunks deleted, so retrieval cannot cite something
|
||||
# that no longer exists.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# AI_ENABLED
|
||||
# Master switch. Fail-closed — must be exactly "true".
|
||||
#
|
||||
# AI_INDEX_DB
|
||||
# SQLite index path. (shipped default: $DATA_DIR/ai_index.db)
|
||||
#
|
||||
# AI_INDEX_BATCH
|
||||
# Chunks per embed request. (shipped default: 32)
|
||||
#
|
||||
# AI_CONNECT_TIMEOUT
|
||||
# Seconds for the reachability probe. (shipped default: 5)
|
||||
#
|
||||
# AI_REQUEST_TIMEOUT
|
||||
# Seconds for a single embed batch. (shipped default: 240)
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_OLLAMA_URL
|
||||
# This host's Ollama endpoint. Empty means no local Ollama.
|
||||
#
|
||||
# HOST*_OLLAMA_EMBED_MODEL
|
||||
# Embedding model. The generation model cannot embed.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ai_index.sh
|
||||
# Incremental refresh — only files whose mtime moved are re-embedded.
|
||||
#
|
||||
# ai_index.sh --force
|
||||
# Full rebuild. Discards the existing index and re-embeds every chunk.
|
||||
#
|
||||
# ai_index.sh --dry-run
|
||||
# Report what would be indexed. Contacts nothing and writes nothing.
|
||||
#
|
||||
# ai_index.sh --status
|
||||
# Show index location, size, chunk counts by kind and section, and last build time.
|
||||
#
|
||||
# ai_index.sh --log
|
||||
# Verbose — per-batch embedding progress.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "${SCRIPT_DIR}/../load_config.sh"
|
||||
|
||||
detect_hosts
|
||||
|
||||
DRY_RUN=false; FORCE=false; STATUS=false; LOG=false
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--dry-run) DRY_RUN=true ;;
|
||||
--force) FORCE=true ;;
|
||||
--status) STATUS=true ;;
|
||||
--log) LOG=true ;;
|
||||
*) echo "Unknown option: $arg" >&2; exit 1 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
CLI="${SCRIPT_DIR}/lib/cli.js"
|
||||
DB="${AI_INDEX_DB:-${DATA_DIR}/ai_index.db}"
|
||||
|
||||
_url_var="${MY_ID}_OLLAMA_URL"
|
||||
_emb_var="${MY_ID}_OLLAMA_EMBED_MODEL"
|
||||
OLLAMA_URL="${!_url_var:-}"
|
||||
EMBED_MODEL="${!_emb_var:-nomic-embed-text}"
|
||||
|
||||
# ── Status ────────────────────────────────────────────────────────────────────────────────────
|
||||
if [[ "$STATUS" == true ]]; then
|
||||
echo "$ICON_GEAR AI Index Status"
|
||||
echo " Enabled: ${AI_ENABLED:-false}"
|
||||
echo " Database: $DB"
|
||||
if [[ -f "$DB" ]]; then
|
||||
echo " Size: $(du -h "$DB" 2>/dev/null | cut -f1)"
|
||||
echo " Chunks: $(sqlite3 "$DB" 'SELECT COUNT(*) FROM vv_chunks;' 2>/dev/null || echo '?')"
|
||||
echo " Files: $(sqlite3 "$DB" 'SELECT COUNT(*) FROM vv_files;' 2>/dev/null || echo '?')"
|
||||
_last=$(sqlite3 "$DB" "SELECT v FROM vv_meta WHERE k='last_index';" 2>/dev/null)
|
||||
[[ -n "$_last" ]] && echo " Last built: $(date -d "@$_last" '+%Y-%m-%d %H:%M:%S' 2>/dev/null)"
|
||||
echo " Model: $(sqlite3 "$DB" "SELECT v FROM vv_meta WHERE k='embed_model';" 2>/dev/null || echo '?')"
|
||||
echo " By kind:"
|
||||
sqlite3 "$DB" "SELECT ' '||kind||': '||COUNT(*) FROM vv_chunks GROUP BY kind ORDER BY COUNT(*) DESC;" 2>/dev/null
|
||||
else
|
||||
echo " Database: not built yet"
|
||||
fi
|
||||
echo " Ollama: ${OLLAMA_URL:-<none local>}"
|
||||
echo " Embed: $EMBED_MODEL"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ── Gate ──────────────────────────────────────────────────────────────────────────────────────
|
||||
if [[ "${AI_ENABLED:-false}" != "true" ]]; then
|
||||
log "AI_ENABLED is not true — skipping index build"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# The mesh shares one AI, and the index belongs to the node that holds the model. A mirror has the
|
||||
# same checkout and could build one, but nothing there would read it: retrieval happens wherever
|
||||
# generation happens, which is the owner.
|
||||
#
|
||||
# A skip, not an error. This is reached from git_pull_execute.sh on every node after every pull;
|
||||
# before the AI became mesh-wide it ran here too and failed on the empty OLLAMA_URL, nightly and
|
||||
# silently, because the caller discards its output.
|
||||
_ai_owner="${AI_OWNER_HOST:-host1}"
|
||||
if [[ "${MY_ID,,}" != "${_ai_owner,,}" ]]; then
|
||||
log "This node is not the AI owner ($_ai_owner) — the index lives there; skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == false && "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
command -v node >/dev/null 2>&1 || { error "node not found — required to build the index"; exit 1; }
|
||||
[[ -f "$CLI" ]] || { error "missing $CLI"; exit 1; }
|
||||
|
||||
# The AI owner has had data/ai since the subsystem was built, so nothing ever created it — cli.js
|
||||
# opens the DB by path and does not make the directory. On a first build the failure surfaces as a
|
||||
# sqlite open error rather than as the missing directory it is.
|
||||
if [[ "$DRY_RUN" == false ]] && ! mkdir -p "$(dirname "$DB")"; then
|
||||
error "Cannot create $(dirname "$DB")"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ -z "$OLLAMA_URL" ]]; then
|
||||
error "${MY_ID}_OLLAMA_URL is empty — no local Ollama to index against"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Pre-flight: fail once, up front, rather than once per batch across the whole corpus.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
if ! curl -sf --max-time "${AI_CONNECT_TIMEOUT:-5}" "${OLLAMA_URL%/}/api/tags" >/dev/null 2>&1; then
|
||||
error "Ollama unreachable at $OLLAMA_URL"
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Build ─────────────────────────────────────────────────────────────────────────────────────
|
||||
_args=(
|
||||
index
|
||||
"--root=${SCRIPTS_DIR}"
|
||||
"--db=${DB}"
|
||||
"--url=${OLLAMA_URL}"
|
||||
"--model=${EMBED_MODEL}"
|
||||
"--batch=${AI_INDEX_BATCH:-32}"
|
||||
"--timeout=$(( ${AI_REQUEST_TIMEOUT:-240} * 1000 ))"
|
||||
)
|
||||
[[ "$FORCE" == true ]] && _args+=(--force)
|
||||
[[ "$DRY_RUN" == true ]] && _args+=(--dry-run)
|
||||
[[ "$LOG" == false ]] && _args+=(--quiet)
|
||||
|
||||
log "$ICON_GEAR Building AI index → $DB"
|
||||
node --no-warnings "$CLI" "${_args[@]}"
|
||||
_rc=$?
|
||||
|
||||
case "$_rc" in
|
||||
0) log "$ICON_DONE AI index build complete" ;;
|
||||
3) warn "AI index built with some batches failed — those files will retry next run" ;;
|
||||
*) error "AI index build failed (exit $_rc)" ;;
|
||||
esac
|
||||
|
||||
exit "$_rc"
|
||||
Executable
+209
@@ -0,0 +1,209 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================== AI Retrieval Query ============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ==============================================================================================
|
||||
# Answers questions about Varaverk from Varaverk's own documentation. Embeds the question,
|
||||
# retrieves the closest chunks from the index AI/ai_index.sh built, and either prints them
|
||||
# directly or passes them to the generation model as grounding context.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
# Two modes over the same retrieval:
|
||||
#
|
||||
# --search print the matching chunks and their sources. No generation model involved,
|
||||
# so it is fast and its output is verbatim repo text.
|
||||
# (default) retrieve, then ask the generation model to answer strictly from what was
|
||||
# retrieved, citing each claim.
|
||||
#
|
||||
# Retrieval is steered by question shape. A question about what prevents something is pushed
|
||||
# toward OPERATIONAL SAFEGUARDS chunks, one about a variable toward CONFIGURATION, one asking
|
||||
# why toward DESIGN PRINCIPLES. This is a score boost, not a filter — a heuristic must not be
|
||||
# able to exclude the chunk that actually holds the answer.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Grounded Or Silent
|
||||
# The prompt instructs the model to answer only from retrieved context and to say what is
|
||||
# missing rather than fill the gap. This repo's conventions are frequently not the
|
||||
# conventional ones, and a confident generic answer about rsync or Docker is worse here
|
||||
# than no answer.
|
||||
#
|
||||
# Sources Are Always Shown
|
||||
# Every answer prints the chunks it drew on. An answer that cannot be traced back to a file
|
||||
# is not usable for changing anything.
|
||||
#
|
||||
# Search Is The Trustworthy Mode
|
||||
# --search returns repo text with nothing generated. When an answer matters, use it.
|
||||
#
|
||||
# Read-Only
|
||||
# Retrieves and answers. Nothing here writes conf, touches state, or runs another script.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Fail-Closed Gate
|
||||
# Exits cleanly unless AI_ENABLED is exactly "true".
|
||||
#
|
||||
# No Root Required
|
||||
# Reads the index and calls Ollama. Nothing it does needs privilege, so it does not ask for
|
||||
# any — this is the one AI script an ordinary user should be able to run.
|
||||
#
|
||||
# Missing Index Is Reported, Not Built
|
||||
# An absent index exits with guidance to run ai_index.sh. Building a corpus-wide index as a
|
||||
# side effect of a question would turn a two-second query into a several-minute one.
|
||||
#
|
||||
# Reachability Pre-flight
|
||||
# Probes Ollama with AI_CONNECT_TIMEOUT before embedding, so an unreachable endpoint fails
|
||||
# immediately with a clear message.
|
||||
#
|
||||
# Bounded Generation
|
||||
# The request is capped at AI_REQUEST_TIMEOUT. A wedged model cannot hang the caller.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# AI_ENABLED
|
||||
# Master switch. Fail-closed — must be exactly "true".
|
||||
#
|
||||
# AI_INDEX_DB
|
||||
# SQLite index to search. (shipped default: $DATA_DIR/ai_index.db)
|
||||
#
|
||||
# AI_SEARCH_K
|
||||
# Chunks retrieved per query. (shipped default: 8)
|
||||
#
|
||||
# AI_SEARCH_PER_FILE
|
||||
# Cap per file, so one document cannot fill the context. (shipped default: 3)
|
||||
#
|
||||
# AI_REQUEST_TIMEOUT
|
||||
# Seconds allowed for generation. (shipped default: 240)
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_OLLAMA_URL / HOST*_OLLAMA_MODEL / HOST*_OLLAMA_EMBED_MODEL
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ai_query.sh "your question"
|
||||
# Retrieve and answer, with sources.
|
||||
#
|
||||
# ai_query.sh --search "your question"
|
||||
# Print matching chunks only. No generation.
|
||||
#
|
||||
# ai_query.sh --section=CONFIGURATION "your question"
|
||||
# Restrict retrieval to one header section.
|
||||
#
|
||||
# ai_query.sh --kind=readme "what is Varaverk"
|
||||
# Restrict retrieval to one chunk origin: header, readme, manual, template, doc.
|
||||
# Use this for definitional and narrative questions. Intent routing boosts header
|
||||
# sections such as PURPOSE, which answers "what does this script do" well but buries
|
||||
# the top-level prose that explains what the system *is*. --kind=readme reaches it.
|
||||
#
|
||||
# ai_query.sh --json "your question"
|
||||
# Machine-readable output for other scripts.
|
||||
#
|
||||
# ai_query.sh --status
|
||||
# Show index and endpoint state.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "${SCRIPT_DIR}/../load_config.sh"
|
||||
|
||||
detect_hosts
|
||||
|
||||
SEARCH_ONLY=false; JSON=false; STATUS=false; SECTION=""; KIND=""; QUERY=""
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--search) SEARCH_ONLY=true ;;
|
||||
--json) JSON=true ;;
|
||||
--status) STATUS=true ;;
|
||||
--section=*) SECTION="${arg#*=}" ;;
|
||||
--kind=*) KIND="${arg#*=}" ;;
|
||||
--*) echo "Unknown option: $arg" >&2; exit 1 ;;
|
||||
*) QUERY="$arg" ;;
|
||||
esac
|
||||
done
|
||||
|
||||
# --kind is a hard filter on chunk origin; --section filters the header section within a
|
||||
# chunk. They answer different questions and compose: --kind=readme --section=PURPOSE is
|
||||
# meaningful. Validated here rather than in the CLI so a typo costs nothing — an unknown kind
|
||||
# silently matches no rows, which reads as "the index has no answer" and is the most
|
||||
# misleading failure this tool can produce.
|
||||
if [[ -n "$KIND" ]]; then
|
||||
case "$KIND" in
|
||||
header|readme|manual|template|doc|ui) ;;
|
||||
*) echo "Unknown --kind=$KIND (expected: header, readme, manual, template, doc, ui)" >&2
|
||||
exit 1 ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
CLI="${SCRIPT_DIR}/lib/cli.js"
|
||||
DB="${AI_INDEX_DB:-${DATA_DIR}/ai_index.db}"
|
||||
|
||||
_url_var="${MY_ID}_OLLAMA_URL"
|
||||
_gen_var="${MY_ID}_OLLAMA_MODEL"
|
||||
_emb_var="${MY_ID}_OLLAMA_EMBED_MODEL"
|
||||
OLLAMA_URL="${!_url_var:-}"
|
||||
GEN_MODEL="${!_gen_var:-}"
|
||||
EMBED_MODEL="${!_emb_var:-nomic-embed-text}"
|
||||
|
||||
if [[ "$STATUS" == true ]]; then
|
||||
echo "$ICON_GEAR AI Query Status"
|
||||
echo " Enabled: ${AI_ENABLED:-false}"
|
||||
echo " Index: $DB $([[ -f "$DB" ]] && echo "($(sqlite3 "$DB" 'SELECT COUNT(*) FROM vv_chunks;' 2>/dev/null) chunks)" || echo '(not built)')"
|
||||
echo " Ollama: ${OLLAMA_URL:-<none>}"
|
||||
echo " Generate: ${GEN_MODEL:-<unset>}"
|
||||
echo " Embed: $EMBED_MODEL"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ "${AI_ENABLED:-false}" != "true" ]]; then
|
||||
echo "AI_ENABLED is not true — AI features are off" >&2
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[[ -z "$QUERY" ]] && { echo "usage: ai_query.sh [--search] [--section=NAME] [--kind=KIND] \"your question\"" >&2; exit 1; }
|
||||
[[ -f "$DB" ]] || { error "No index at $DB — run AI/ai_index.sh first"; exit 1; }
|
||||
[[ -f "$CLI" ]] || { error "missing $CLI"; exit 1; }
|
||||
command -v node >/dev/null 2>&1 || { error "node not found"; exit 1; }
|
||||
[[ -z "$OLLAMA_URL" ]] && { error "${MY_ID}_OLLAMA_URL is empty"; exit 1; }
|
||||
|
||||
if ! curl -sf --max-time "${AI_CONNECT_TIMEOUT:-5}" "${OLLAMA_URL%/}/api/tags" >/dev/null 2>&1; then
|
||||
error "Ollama unreachable at $OLLAMA_URL"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
_args=("--db=${DB}" "--url=${OLLAMA_URL}" "--query=${QUERY}")
|
||||
[[ -n "$SECTION" ]] && _args+=("--section=${SECTION}")
|
||||
[[ -n "$KIND" ]] && _args+=("--kind=${KIND}")
|
||||
[[ "$JSON" == true ]] && _args+=(--json)
|
||||
|
||||
if [[ "$SEARCH_ONLY" == true ]]; then
|
||||
node --no-warnings "$CLI" search "${_args[@]}" \
|
||||
"--model=${EMBED_MODEL}" \
|
||||
"--k=${AI_SEARCH_K:-8}" "--per-file=${AI_SEARCH_PER_FILE:-3}"
|
||||
else
|
||||
[[ -z "$GEN_MODEL" ]] && { error "${MY_ID}_OLLAMA_MODEL is empty — needed for generation"; exit 1; }
|
||||
# Token accounting. Passed in rather than re-read in node, so the conf stays the shell's job
|
||||
# and cli.js keeps taking everything it needs as arguments. Omitting either flag simply
|
||||
# skips the row — the CLI must still work when called by hand outside this wrapper.
|
||||
node --no-warnings "$CLI" ask "${_args[@]}" \
|
||||
"--model=${GEN_MODEL}" "--embed-model=${EMBED_MODEL}" \
|
||||
"--k=${AI_SEARCH_K:-8}" "--per-file=${AI_SEARCH_PER_FILE:-3}" \
|
||||
"--timeout=$(( ${AI_REQUEST_TIMEOUT:-240} * 1000 ))" \
|
||||
"--token-db=${AI_TOKEN_DB:-}" "--token-host=$(echo "$MY_ID" | tr '[:upper:]' '[:lower:]')"
|
||||
fi
|
||||
Executable
+236
@@ -0,0 +1,236 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================= AI Token Ledger Sync ===========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Pulls each partner's AI token ledger into a RAM cache at /tmp/.cache/vv/ai/, so the AI tab
|
||||
# can report usage for the whole fleet instead of only the host the browser happens to be on.
|
||||
#
|
||||
# Each host records its own turns to data/ai_token_history.db and nothing syncs that file, so
|
||||
# without this a host can only ever see its own totals. The tab is careful to say "not collected
|
||||
# here" rather than 0 for a partner it cannot see; this script is what turns that into a number.
|
||||
#
|
||||
# Same trick as conf_sync.sh, and deliberately so — resolve the partner over Tailscale, scp one
|
||||
# small file into a tmpfs cache, let the reader treat a missing file as "unknown".
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Gates — PARTNERSHIP_ENABLED, AI_ENABLED, AI_TOKEN_SYNC_ENABLED
|
||||
# 2. Per partner:
|
||||
# a. Resolve their Tailscale IP
|
||||
# b. Resolve their SCRIPTS_DIR from their varaverk.cfg (they may be in appdata mode)
|
||||
# c. scp their data/ai_token_history.db → $AI_TOKEN_CACHE_DIR/<slot>.tokens.db
|
||||
#
|
||||
# Pull only, no push. conf_sync.sh pushes as well because a partner may be unable to reach us
|
||||
# and still needs our credentials; nothing here is needed by anyone else, and a reader that
|
||||
# fetches its own data controls its own freshness rather than depending on the partner's cron.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# A missing file means unknown, never zero.
|
||||
# The whole point of the tab reporting "not collected here" is that it is a different claim
|
||||
# from "this partner spent nothing". If a partner is dark, unreachable or has never run a turn,
|
||||
# there is simply no cache file, and the reader is expected to say so rather than render a 0
|
||||
# that looks like a measurement.
|
||||
#
|
||||
# The reader pulls; nobody pushes.
|
||||
# conf_sync.sh pushes as well, because a partner that cannot reach us still needs our
|
||||
# credentials. Nothing here is needed by anyone else, so a host that wants fleet totals fetches
|
||||
# them and owns its own freshness instead of depending on someone else's cron having run.
|
||||
#
|
||||
# RAM, not flash.
|
||||
# The cache lands in tmpfs. It is a copy of a file that already exists on the partner and is
|
||||
# rebuilt on the next pass, so writing it to flash would cost wear for something that is never
|
||||
# worth surviving a reboot.
|
||||
#
|
||||
# Same shape as conf_sync.sh, deliberately.
|
||||
# Resolve over Tailscale, scp one small file into a tmpfs cache, let a missing file mean
|
||||
# unknown. A second transport pattern for a second small file would be a second set of
|
||||
# failure modes to learn.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# An unreachable partner is not a failure.
|
||||
# HOST2 is expected to be down for long stretches during onboarding. A warn every four
|
||||
# hours would train the operator to ignore this script's output, and the AI diagnostic
|
||||
# path treats every log WARN as actionable. Unresolvable partners are counted and
|
||||
# reported once at info level; only a partner that resolves and then fails to transfer
|
||||
# is treated as an error.
|
||||
#
|
||||
# The cache is never written directly.
|
||||
# scp lands on a .part file that is renamed into place, so a transfer interrupted halfway
|
||||
# cannot leave the reader parsing half a ledger. A truncated final row would be skipped by
|
||||
# the field-count check on the PHP side, but a torn file should not reach it at all.
|
||||
#
|
||||
# Nothing is ever written back to the partner.
|
||||
# This script only reads. A bug here cannot corrupt a partner's accounting.
|
||||
#
|
||||
# The cache is tmpfs and deliberately not preserved.
|
||||
# Unlike the conf cache there is no save/restore pair. Stale counters are worse than
|
||||
# absent ones: absent reads as "not collected here", stale reads as fact.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# AI_ENABLED Whole AI subsystem gate
|
||||
# AI_TOKEN_SYNC_ENABLED This script's own toggle (default: true)
|
||||
# PARTNERSHIP_ENABLED Checked via require_partnership()
|
||||
# SSH_KEY Key used for all partner ssh/scp operations
|
||||
#
|
||||
# load_config.sh
|
||||
#
|
||||
# AI_TOKEN_CACHE_DIR tmpfs directory the tab reads partner ledgers from
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST* — hostnames used to build the partner list via detect_hosts()
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ai_token_sync.sh Pull every reachable partner's ledger
|
||||
# ai_token_sync.sh --dry-run Report what would be pulled, transfer nothing
|
||||
# ai_token_sync.sh --log Verbose output
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
require_partnership
|
||||
|
||||
if [[ "${AI_ENABLED:-false}" != true ]]; then
|
||||
log "AI_ENABLED=false — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ "${AI_TOKEN_SYNC_ENABLED:-true}" == false ]]; then
|
||||
log "AI_TOKEN_SYNC_ENABLED=false — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
CACHE_DIR="$AI_TOKEN_CACHE_DIR"
|
||||
SSH_TIMEOUT=10
|
||||
|
||||
# Mirrors conf_sync.sh — the remote may be in appdata storage mode, so its ledger is not
|
||||
# necessarily under /boot.
|
||||
_remote_scripts_dir() {
|
||||
local ip="$1"
|
||||
local cfg line sd
|
||||
cfg=$(timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
|
||||
"root@${ip}" "cat /boot/config/plugins/varaverk/varaverk.cfg 2>/dev/null" 2>/dev/null) || true
|
||||
while IFS= read -r line; do
|
||||
[[ "$line" == SCRIPTS_DIR=* ]] || continue
|
||||
sd="${line#SCRIPTS_DIR=}"; sd="${sd//\"/}"; sd="${sd//\'/}"
|
||||
echo "$sd"; return
|
||||
done <<< "$cfg"
|
||||
echo "/boot/config/plugins/varaverk"
|
||||
}
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
||||
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
mkdir -p "$CACHE_DIR" && chmod 755 "$CACHE_DIR"
|
||||
fi
|
||||
|
||||
PULLED=0
|
||||
OFFLINE=0
|
||||
FAILED=0
|
||||
|
||||
for host_var in $(compgen -v | grep -E '^HOST[0-9]+$' | sort); do
|
||||
partner_host="${!host_var}"
|
||||
[[ -z "$partner_host" ]] && continue
|
||||
[[ "${host_var,,}" == "${MY_ID,,}" ]] && continue
|
||||
|
||||
partner_slot="${host_var,,}"
|
||||
partner_ip=$(resolve_tailscale_ip "$partner_host" 2>/dev/null || true)
|
||||
|
||||
if [[ -z "$partner_ip" ]]; then
|
||||
log "$partner_host — unresolvable, leaving any cached ledger as-is"
|
||||
(( OFFLINE++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Liveness and ledger presence are separate questions, probed separately on purpose.
|
||||
# Tailscale hands back an IP for a peer that is registered but powered off, so resolving
|
||||
# proves nothing. And a single `ssh test -f` answers both questions at once: it fails
|
||||
# identically whether the host is down or the file is simply absent. Treating that one
|
||||
# failure as "no ledger" would delete a perfectly good cached copy every time the partner
|
||||
# blinked — turning "synced 3h ago" into "not collected here" on a transient.
|
||||
if ! timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
|
||||
"root@${partner_ip}" true 2>/dev/null; then
|
||||
log "$partner_host — not answering, keeping any cached ledger as-is"
|
||||
(( OFFLINE++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
remote_sd=$(_remote_scripts_dir "$partner_ip")
|
||||
remote_db="${remote_sd}/data/ai_token_history.db"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would pull $partner_host:$remote_db → $CACHE_DIR/${partner_slot}.tokens.db"
|
||||
continue
|
||||
fi
|
||||
|
||||
# Reachable, but nothing recorded there — a partner with AI off, or one that has simply
|
||||
# never been asked anything. Now that liveness is established this is a real answer, so the
|
||||
# stale copy goes: the tab should say "not collected here", not quote a number from before.
|
||||
if ! timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
|
||||
"root@${partner_ip}" "[[ -f '$remote_db' ]]" 2>/dev/null; then
|
||||
log "$partner_host — reachable, but no ledger there yet"
|
||||
rm -f "$CACHE_DIR/${partner_slot}.tokens.db"
|
||||
(( OFFLINE++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
if timeout "$SSH_TIMEOUT" scp -i "$SSH_KEY" \
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" -o BatchMode=yes -o StrictHostKeyChecking=no \
|
||||
"root@${partner_ip}:${remote_db}" \
|
||||
"$CACHE_DIR/${partner_slot}.tokens.db.part" 2>/dev/null \
|
||||
&& mv -f "$CACHE_DIR/${partner_slot}.tokens.db.part" "$CACHE_DIR/${partner_slot}.tokens.db"; then
|
||||
chmod 644 "$CACHE_DIR/${partner_slot}.tokens.db" 2>/dev/null
|
||||
echo "Pulled ${partner_slot} ledger from $partner_host ✅"
|
||||
(( PULLED++ ))
|
||||
else
|
||||
rm -f "$CACHE_DIR/${partner_slot}.tokens.db.part"
|
||||
warn "Could not pull ${partner_slot} ledger from $partner_host"
|
||||
(( FAILED++ ))
|
||||
fi
|
||||
done
|
||||
|
||||
# Zero counts are omitted rather than printed. ${VAR:+...} keeps "0" because it is a non-empty
|
||||
# string, and a summary that always ends "0 failed" is what teaches you to stop reading it.
|
||||
_summary="AI token sync complete — pulled $PULLED"
|
||||
(( OFFLINE > 0 )) && _summary+=", $OFFLINE unavailable"
|
||||
(( FAILED > 0 )) && _summary+=", $FAILED failed"
|
||||
info "$_summary"
|
||||
|
||||
# Only a partner that answered and then failed the transfer is worth an exit code. An absent
|
||||
# partner is the normal state whenever a partner is not yet onboarded.
|
||||
[[ "$FAILED" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
+276
@@ -0,0 +1,276 @@
|
||||
'use strict';
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
// Chunker — turns repo files into retrieval units.
|
||||
//
|
||||
// The whole point of the header audit is that chunk boundaries are deterministic here. Bash
|
||||
// scripts split on their six section names, markdown on its headings, conf templates on their
|
||||
// ━━━ section rules. Nothing is split on a fixed token window, so no chunk ever contains half
|
||||
// of one idea and half of another.
|
||||
//
|
||||
// Every chunk carries its section name as its own field, because that is the metadata that
|
||||
// lets retrieval filter by question shape before it ever computes similarity.
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
const BASH_SECTIONS = [
|
||||
'PURPOSE', 'OPERATIONAL MODEL', 'DESIGN PRINCIPLES',
|
||||
'OPERATIONAL SAFEGUARDS', 'CONFIGURATION', 'RUNTIME MODES',
|
||||
];
|
||||
|
||||
// PHP headers reuse the first three names deliberately, then diverge per layer.
|
||||
const PHP_SECTIONS = [
|
||||
'PURPOSE', 'OPERATIONAL MODEL', 'DESIGN PRINCIPLES', 'OPERATIONAL SAFEGUARDS',
|
||||
'STATUS', 'EXPORTS', 'REQUEST CONTRACT', 'SIDE EFFECTS', 'RENDERS', 'DEPENDS ON',
|
||||
'CONFIGURATION',
|
||||
];
|
||||
|
||||
const MIN_CHARS = 40; // below this a chunk carries no retrievable meaning
|
||||
const MAX_CHARS = 6000; // above this, split on blank lines — protects the embed window
|
||||
|
||||
// Banner rules and box-drawing art are everywhere in this repo's headers. They carry no
|
||||
// meaning to embed, and a chunk that is mostly rule characters is pure noise in the index.
|
||||
// Measure a chunk by what is left after the decoration is removed, not by raw length.
|
||||
function meaningful(s) {
|
||||
return s.replace(/[═─━=_#\/*\s|+.-]/g, '').length;
|
||||
}
|
||||
const MIN_MEANINGFUL = 30;
|
||||
|
||||
function isRealHeading(h) {
|
||||
return !!h && /[A-Za-z0-9]/.test(h.replace(/[═─━=_]/g, ''));
|
||||
}
|
||||
|
||||
function stripPrefix(line, prefix) {
|
||||
// '# text' -> 'text' '// text' -> 'text'
|
||||
const re = new RegExp('^\\s*' + prefix + '\\s?');
|
||||
return line.replace(re, '');
|
||||
}
|
||||
|
||||
// ── Comment-header sectioning, shared by bash (#) and PHP (//) ────────────────────────────────
|
||||
function sectionsFromCommentHeader(text, prefix, names) {
|
||||
const lines = text.split('\n');
|
||||
const nameSet = new Set(names);
|
||||
const found = [];
|
||||
|
||||
// headerEnd matters as much as the section starts. The last section (RUNTIME MODES in bash,
|
||||
// DEPENDS ON in a page) would otherwise run to EOF and sweep up every unrelated comment in
|
||||
// the file — scheduler.php alone contributed an 11k-char chunk of unrelated inline comments.
|
||||
let headerEnd = lines.length;
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const raw = lines[i];
|
||||
if (!new RegExp('^\\s*' + prefix).test(raw)) {
|
||||
// Header block ends at the first non-comment, non-blank line past the shebang.
|
||||
// '<?php' and '?>' bracket a PHP header block and are not the end of it.
|
||||
const t = raw.trim();
|
||||
if (t !== '' && !/^#!/.test(t) && t !== '<?php' && t !== '?>' && found.length) {
|
||||
headerEnd = i;
|
||||
break;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
const inner = stripPrefix(raw, prefix).trim();
|
||||
if (nameSet.has(inner)) found.push({ name: inner, start: i });
|
||||
}
|
||||
|
||||
const out = [];
|
||||
for (let k = 0; k < found.length; k++) {
|
||||
const start = found[k].start + 1;
|
||||
const end = k + 1 < found.length ? found[k + 1].start : headerEnd;
|
||||
const body = lines.slice(start, end)
|
||||
.filter(l => new RegExp('^\\s*' + prefix).test(l))
|
||||
.map(l => stripPrefix(l, prefix))
|
||||
// drop pure separator rules (════, ────, ━━━) — they carry no meaning
|
||||
.filter(l => !/^[\s═─━=_-]*$/.test(l) || l.trim() === '')
|
||||
.join('\n')
|
||||
.replace(/\n{3,}/g, '\n\n')
|
||||
.trim();
|
||||
if (meaningful(body) >= MIN_MEANINGFUL) {
|
||||
for (const p of splitNamedParagraphs(body))
|
||||
out.push({ section: found[k].name, title: p.title, content: p.content });
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ── Named-paragraph sub-chunking ──────────────────────────────────────────────────────────────
|
||||
// The header convention writes safeguards and principles as named paragraphs: an unindented
|
||||
// title line followed by an indented body. Embedding a whole section as one unit dilutes them —
|
||||
// rsync.sh's OPERATIONAL SAFEGUARDS holds eight distinct guarantees in 2.8k chars, and a query
|
||||
// about one of them scored below unrelated chunks because the other seven dominated the vector.
|
||||
// Splitting on the title lines is what makes a specific question find a specific answer.
|
||||
//
|
||||
// The section name is carried onto every sub-chunk, so section routing still works; the
|
||||
// paragraph title becomes the chunk's heading.
|
||||
function splitNamedParagraphs(body) {
|
||||
const lines = body.split('\n');
|
||||
const marks = [];
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const l = lines[i];
|
||||
if (!l.trim()) continue;
|
||||
if (/^\s/.test(l)) continue; // indented => body, not a title
|
||||
if (/^[-*•]/.test(l.trim())) continue; // list item, not a title
|
||||
if (l.trim().length > 80) continue; // a long line is prose, not a heading
|
||||
if (/[.:;,]$/.test(l.trim())) continue; // ends like a sentence
|
||||
|
||||
// The decisive signal: a real title is followed by an indented body. Wrapped prose is
|
||||
// followed by more unindented prose. Without this check, any short line in a paragraph
|
||||
// that happened to wrap became a spurious chunk boundary mid-sentence.
|
||||
let j = i + 1;
|
||||
while (j < lines.length && !lines[j].trim()) j++;
|
||||
if (j >= lines.length || !/^\s+\S/.test(lines[j])) continue;
|
||||
|
||||
marks.push(i);
|
||||
}
|
||||
// Fewer than two titles means this section is not written as named paragraphs — keep it whole.
|
||||
if (marks.length < 2) return [{ title: null, content: body }];
|
||||
|
||||
const out = [];
|
||||
if (marks[0] > 0) {
|
||||
const pre = lines.slice(0, marks[0]).join('\n').trim();
|
||||
if (meaningful(pre) >= MIN_MEANINGFUL) out.push({ title: null, content: pre });
|
||||
}
|
||||
for (let k = 0; k < marks.length; k++) {
|
||||
const start = marks[k];
|
||||
const end = k + 1 < marks.length ? marks[k + 1] : lines.length;
|
||||
const title = lines[start].trim();
|
||||
const content = lines.slice(start, end).join('\n').trim();
|
||||
if (meaningful(content) >= MIN_MEANINGFUL) out.push({ title, content });
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ── Markdown: split on ## headings, keep the heading with its body ─────────────────────────────
|
||||
function sectionsFromMarkdown(text) {
|
||||
const lines = text.split('\n');
|
||||
const marks = [];
|
||||
let fence = false;
|
||||
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
if (/^\s*```/.test(lines[i])) { fence = !fence; continue; }
|
||||
if (fence) continue;
|
||||
// A heading whose text is nothing but rule characters is a banner, not a section. The
|
||||
// house style opens a document with three of them —
|
||||
// # ━━━━━━━━
|
||||
// # 🏠 VARAVERK
|
||||
// # ━━━━━━━━
|
||||
// — and treating each as a boundary split the title into a section of its own, too small
|
||||
// to survive, then gave the paragraph that actually defines the project a chunk headed by
|
||||
// the rule beneath it: no heading, and content opening with 75 identical glyphs. That is
|
||||
// why "what is Varaverk" returned five script PURPOSE headers and never README.md, which
|
||||
// has been indexed the whole time. 35 of the 40 markdown files here open this way.
|
||||
//
|
||||
// Whole-string test, so an ordinary heading containing a dash is unaffected — "Set-up"
|
||||
// does not reduce to empty, and a line of dashes does.
|
||||
if (!/^#{1,3}\s+\S/.test(lines[i])) continue;
|
||||
if (lines[i].replace(/^#+\s*/, '').replace(/[━─═=~_*\-\s]+/gu, '') === '') continue;
|
||||
marks.push(i);
|
||||
}
|
||||
if (!marks.length) return [{ heading: null, content: text.trim() }];
|
||||
|
||||
const out = [];
|
||||
// preamble before the first heading
|
||||
if (marks[0] > 0) {
|
||||
const pre = lines.slice(0, marks[0]).join('\n').trim();
|
||||
if (meaningful(pre) >= MIN_MEANINGFUL) out.push({ heading: null, content: pre });
|
||||
}
|
||||
for (let k = 0; k < marks.length; k++) {
|
||||
const start = marks[k];
|
||||
const end = k + 1 < marks.length ? marks[k + 1] : lines.length;
|
||||
const heading = lines[start].replace(/^#+\s*/, '').replace(/[━─═]+/g, '').trim();
|
||||
const content = lines.slice(start, end).join('\n').trim();
|
||||
if (meaningful(content) >= MIN_MEANINGFUL) out.push({ heading: isRealHeading(heading) ? heading : null, content });
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// ── Conf templates: split on the ━━━ / ── section rules ───────────────────────────────────────
|
||||
function sectionsFromConfTemplate(text) {
|
||||
const lines = text.split('\n');
|
||||
const marks = [];
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
const m = lines[i].match(/^#\s*[━─]{2,}\s*(.+?)\s*[━─]{2,}\s*$/);
|
||||
if (m && isRealHeading(m[1])) marks.push({ i, name: m[1].trim() });
|
||||
}
|
||||
if (!marks.length) return [];
|
||||
|
||||
const out = [];
|
||||
for (let k = 0; k < marks.length; k++) {
|
||||
const start = marks[k].i;
|
||||
const end = k + 1 < marks.length ? marks[k + 1].i : lines.length;
|
||||
const content = lines.slice(start, end).join('\n').replace(/\n{3,}/g, '\n\n').trim();
|
||||
if (meaningful(content) >= MIN_MEANINGFUL) out.push({ heading: marks[k].name, content });
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
// Oversized chunks split on blank lines rather than mid-sentence.
|
||||
function capSize(chunks) {
|
||||
const out = [];
|
||||
for (const c of chunks) {
|
||||
if (c.content.length <= MAX_CHARS) { out.push(c); continue; }
|
||||
const paras = c.content.split(/\n\s*\n/);
|
||||
let buf = [], len = 0, part = 1;
|
||||
const flush = () => {
|
||||
if (!buf.length) return;
|
||||
out.push({ ...c, content: buf.join('\n\n'), part: part++ });
|
||||
buf = []; len = 0;
|
||||
};
|
||||
for (const p of paras) {
|
||||
if (len + p.length > MAX_CHARS && buf.length) flush();
|
||||
buf.push(p); len += p.length + 2;
|
||||
}
|
||||
flush();
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
function classify(rel) {
|
||||
const base = path.basename(rel);
|
||||
if (rel.startsWith('Deployment/') && rel.endsWith('.template')) return 'template';
|
||||
// WebGUI page docs, written for whoever is using the tab rather than maintaining it. Their
|
||||
// own kind because every other kind here answers a maintainer's question: an operator asking
|
||||
// "how do I stop this" needs the click path, and a corpus that is three-quarters script
|
||||
// headers will otherwise always answer in conf edits. Matched on the folder, not the
|
||||
// filename, so these can be named whatever reads best.
|
||||
if (rel.startsWith('Plugin/unraid/pages/readme/') && base.endsWith('.md')) return 'ui';
|
||||
if (base.endsWith('.md')) {
|
||||
if (base.startsWith('Manual')) return 'manual';
|
||||
if (base.startsWith('README') || base === 'README.md') return 'readme';
|
||||
return 'doc';
|
||||
}
|
||||
if (base.endsWith('.sh')) return 'header';
|
||||
if (base.endsWith('.php')) return 'header';
|
||||
return 'other';
|
||||
}
|
||||
|
||||
function chunkFile(absPath, rel) {
|
||||
const text = fs.readFileSync(absPath, 'utf8');
|
||||
const kind = classify(rel);
|
||||
let raw = [];
|
||||
|
||||
if (kind === 'header' && rel.endsWith('.sh')) {
|
||||
raw = sectionsFromCommentHeader(text, '#', BASH_SECTIONS)
|
||||
.map(s => ({ section: s.section, heading: s.title || null, content: s.content }));
|
||||
} else if (kind === 'header' && rel.endsWith('.php')) {
|
||||
raw = sectionsFromCommentHeader(text, '//', PHP_SECTIONS)
|
||||
.map(s => ({ section: s.section, heading: s.title || null, content: s.content }));
|
||||
} else if (kind === 'template') {
|
||||
raw = sectionsFromConfTemplate(text)
|
||||
.map(s => ({ section: null, heading: s.heading, content: s.content }));
|
||||
} else if (kind === 'readme' || kind === 'manual' || kind === 'doc' || kind === 'ui') {
|
||||
raw = sectionsFromMarkdown(text)
|
||||
.map(s => ({ section: null, heading: s.heading, content: s.content }));
|
||||
}
|
||||
|
||||
return capSize(raw).map(c => ({
|
||||
path: rel,
|
||||
kind,
|
||||
section: c.section || null,
|
||||
heading: c.heading || null,
|
||||
part: c.part || null,
|
||||
content: c.content,
|
||||
}));
|
||||
}
|
||||
|
||||
module.exports = { chunkFile, classify, BASH_SECTIONS, PHP_SECTIONS };
|
||||
+209
@@ -0,0 +1,209 @@
|
||||
'use strict';
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
// CLI bridge — the thin layer the bash entry points call.
|
||||
//
|
||||
// The bash scripts own configuration, gating, locking and logging, exactly as they do for every
|
||||
// other Varaverk job. This file owns only the work that is genuinely awkward in bash: float
|
||||
// vector math and SQLite BLOBs. That split follows the existing api_cache_writer.sh precedent —
|
||||
// a bash shim in front of the language that fits the task.
|
||||
//
|
||||
// Every value arrives as an argument or an environment variable read by the caller. This file
|
||||
// never reads a conf file itself, so there is exactly one place that decides what the settings
|
||||
// are.
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
const fs = require('fs');
|
||||
const { buildIndex } = require('./index.js');
|
||||
const { search } = require('./search.js');
|
||||
|
||||
function arg(name, dflt) {
|
||||
const p = `--${name}=`;
|
||||
const hit = process.argv.find(a => a.startsWith(p));
|
||||
return hit ? hit.slice(p.length) : dflt;
|
||||
}
|
||||
function flag(name) {
|
||||
return process.argv.includes(`--${name}`);
|
||||
}
|
||||
|
||||
// Token accounting. Writes the same row shape as the WebGUI worker into the same file — one
|
||||
// ledger for both paths, or the totals quietly come to mean "whatever the tab happened to do".
|
||||
// Skipped silently when the caller passes neither flag, because cli.js has to stay runnable by
|
||||
// hand. Best-effort: a failed append must never cost a caller an answer it already has.
|
||||
//
|
||||
// Trimming is deliberately not done here. The PHP side prunes on write, and duplicating a
|
||||
// read-modify-write of the whole file in a second language is how the two drift apart.
|
||||
function recordTokens(profile, prompt, completion, tokS) {
|
||||
const db = arg('token-db', ''), host = arg('token-host', '');
|
||||
if (!db || !host || (prompt <= 0 && completion <= 0)) return;
|
||||
const d = new Date();
|
||||
const p2 = n => String(n).padStart(2, '0');
|
||||
const row = [
|
||||
`${d.getFullYear()}-${p2(d.getMonth() + 1)}-${p2(d.getDate())}`,
|
||||
`${p2(d.getHours())}:${p2(d.getMinutes())}:${p2(d.getSeconds())}`,
|
||||
host, profile, 'cli', prompt, completion,
|
||||
tokS === null ? '' : tokS.toFixed(1),
|
||||
].join('|') + '\n';
|
||||
try { fs.appendFileSync(db, row); } catch { /* accounting is not the answer */ }
|
||||
}
|
||||
|
||||
function fail(msg, code = 1) {
|
||||
console.error(msg);
|
||||
process.exit(code);
|
||||
}
|
||||
|
||||
async function cmdIndex() {
|
||||
const root = arg('root');
|
||||
const db = arg('db');
|
||||
const url = arg('url');
|
||||
const model = arg('model', 'nomic-embed-text');
|
||||
if (!root || !db || !url) fail('index: --root, --db and --url are required');
|
||||
|
||||
const quiet = flag('quiet');
|
||||
let stats;
|
||||
try {
|
||||
stats = await buildIndex({
|
||||
root, dbPath: db, url, model,
|
||||
batch: parseInt(arg('batch', '32'), 10),
|
||||
timeout: parseInt(arg('timeout', '120000'), 10),
|
||||
force: flag('force'),
|
||||
dryRun: flag('dry-run'),
|
||||
onProgress: p => {
|
||||
if (p.error) console.error(`embed batch failed: ${p.error}`);
|
||||
else if (!quiet && p.done % 320 === 0) console.log(` embedded ${p.done}/${p.total}`);
|
||||
},
|
||||
});
|
||||
} catch (e) {
|
||||
fail(`index failed: ${e.message}`, 2);
|
||||
}
|
||||
|
||||
if (flag('json')) { console.log(JSON.stringify(stats)); return; }
|
||||
if (stats.dryRun) {
|
||||
console.log(`DRY RUN — ${stats.files} file(s) would be indexed, ${stats.chunks} chunk(s) embedded`);
|
||||
console.log(` ${stats.skipped} unchanged, ${stats.removed} stale entr(ies) would be dropped`);
|
||||
return;
|
||||
}
|
||||
console.log(`indexed ${stats.files} file(s), ${stats.chunks} chunk(s) embedded`);
|
||||
console.log(` ${stats.skipped} unchanged, ${stats.removed} removed, ${stats.failed} failed`);
|
||||
console.log(` index now holds ${stats.total} chunk(s)`);
|
||||
// A partial index is usable but not complete — say so in the exit code so a caller can act.
|
||||
if (stats.failed) process.exit(3);
|
||||
}
|
||||
|
||||
async function cmdSearch() {
|
||||
const db = arg('db');
|
||||
const url = arg('url');
|
||||
const model = arg('model', 'nomic-embed-text');
|
||||
const q = arg('query');
|
||||
if (!db || !url || !q) fail('search: --db, --url and --query are required');
|
||||
|
||||
let r;
|
||||
try {
|
||||
r = await search({
|
||||
dbPath: db, url, embedModel: model, query: q,
|
||||
k: parseInt(arg('k', '8'), 10),
|
||||
perFile: parseInt(arg('per-file', '3'), 10),
|
||||
section: arg('section', null),
|
||||
kind: arg('kind', null),
|
||||
});
|
||||
} catch (e) {
|
||||
fail(`search failed: ${e.message}`, 2);
|
||||
}
|
||||
|
||||
if (flag('json')) { console.log(JSON.stringify(r)); return; }
|
||||
if (!r.results.length) { console.log('no matches'); return; }
|
||||
if (r.intents.length) console.log(`intent: ${r.intents.join(', ')}\n`);
|
||||
for (const x of r.results) {
|
||||
const label = x.heading || x.section || '-';
|
||||
console.log(`── ${x.score.toFixed(3)} ${x.path} [${x.section || x.kind}] ${label}`);
|
||||
console.log(x.content.split('\n').map(l => ' ' + l).join('\n'));
|
||||
console.log('');
|
||||
}
|
||||
}
|
||||
|
||||
// Retrieval + generation. The prompt is built here so the context block and the instructions
|
||||
// stay in one reviewable place.
|
||||
async function cmdAsk() {
|
||||
const db = arg('db');
|
||||
const url = arg('url');
|
||||
const embed = arg('embed-model', 'nomic-embed-text');
|
||||
const gen = arg('model');
|
||||
const q = arg('query');
|
||||
const timeout = parseInt(arg('timeout', '240000'), 10);
|
||||
if (!db || !url || !gen || !q) fail('ask: --db, --url, --model and --query are required');
|
||||
|
||||
let r;
|
||||
try {
|
||||
// section and kind must be forwarded here too. They were not, so both filters worked
|
||||
// under `search` and were silently ignored under `ask` — the documented
|
||||
// --section=CONFIGURATION usage retrieved from the whole index and the answer looked
|
||||
// plausible, which is the worst way for a filter to fail.
|
||||
r = await search({
|
||||
dbPath: db, url, embedModel: embed, query: q,
|
||||
k: parseInt(arg('k', '6'), 10), perFile: parseInt(arg('per-file', '2'), 10),
|
||||
section: arg('section', null),
|
||||
kind: arg('kind', null),
|
||||
});
|
||||
} catch (e) {
|
||||
fail(`retrieval failed: ${e.message}`, 2);
|
||||
}
|
||||
if (!r.results.length) fail('no relevant context found in the index', 4);
|
||||
|
||||
const context = r.results.map((x, i) => {
|
||||
const label = [x.path, x.section, x.heading].filter(Boolean).join(' › ');
|
||||
return `[${i + 1}] ${label}\n${x.content}`;
|
||||
}).join('\n\n');
|
||||
|
||||
const prompt =
|
||||
`You are answering questions about Varaverk, a two-server self-healing home media ecosystem.
|
||||
|
||||
Answer ONLY from the context below. If the context does not contain the answer, say so plainly
|
||||
and name what is missing — do not fill the gap from general knowledge about Linux, Docker or
|
||||
rsync, because this system's conventions are frequently not the conventional ones.
|
||||
|
||||
Cite the source of each claim as [n]. Be concise and concrete.
|
||||
|
||||
CONTEXT
|
||||
${context}
|
||||
|
||||
QUESTION
|
||||
${q}
|
||||
|
||||
ANSWER`;
|
||||
|
||||
let res;
|
||||
try {
|
||||
res = await fetch(`${url.replace(/\/$/, '')}/api/generate`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
model: gen, prompt, stream: false,
|
||||
options: { temperature: 0.2, num_ctx: 8192 },
|
||||
}),
|
||||
signal: AbortSignal.timeout(timeout),
|
||||
});
|
||||
} catch (e) {
|
||||
fail(`generation failed: ${e.message}`, 2);
|
||||
}
|
||||
if (!res.ok) fail(`generation HTTP ${res.status}`, 2);
|
||||
const j = await res.json();
|
||||
|
||||
// 'varaverk' rather than a CLI-specific name: this path retrieves and cites, so it is the
|
||||
// same kind of turn the tab's default profile runs, and the two should aggregate together.
|
||||
recordTokens('varaverk', j.prompt_eval_count || 0, j.eval_count || 0,
|
||||
j.eval_duration > 0 ? (j.eval_count / (j.eval_duration / 1e9)) : null);
|
||||
|
||||
if (flag('json')) {
|
||||
console.log(JSON.stringify({ answer: j.response, sources: r.results.map(x => ({ path: x.path, section: x.section, heading: x.heading, score: x.score })) }));
|
||||
return;
|
||||
}
|
||||
console.log((j.response || '').trim());
|
||||
console.log('\nSources:');
|
||||
r.results.forEach((x, i) => {
|
||||
console.log(` [${i + 1}] ${[x.path, x.section, x.heading].filter(Boolean).join(' › ')}`);
|
||||
});
|
||||
}
|
||||
|
||||
const cmd = process.argv[2];
|
||||
const table = { index: cmdIndex, search: cmdSearch, ask: cmdAsk };
|
||||
if (!table[cmd]) fail(`usage: cli.js <index|search|ask> [--flags]`);
|
||||
table[cmd]().catch(e => fail(e.message, 2));
|
||||
+201
@@ -0,0 +1,201 @@
|
||||
'use strict';
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
// Indexer — chunk the repo, embed each chunk, store vectors in SQLite.
|
||||
//
|
||||
// Incremental by file mtime: a file whose mtime has not moved since its last index is skipped
|
||||
// entirely, so a routine re-index costs seconds rather than re-embedding the whole corpus.
|
||||
//
|
||||
// Vectors are stored as raw little-endian float32 BLOBs. nomic-embed-text returns L2-normalised
|
||||
// vectors, so cosine similarity is a plain dot product at query time — no normalising, no
|
||||
// magnitude cache. PHP can read the same blobs with unpack('f*', $blob) when the UI needs them.
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const { execSync } = require('child_process');
|
||||
const { DatabaseSync } = require('node:sqlite');
|
||||
const { chunkFile, classify } = require('./chunk.js');
|
||||
|
||||
const SCHEMA = `
|
||||
CREATE TABLE IF NOT EXISTS vv_files (
|
||||
path TEXT PRIMARY KEY,
|
||||
mtime INTEGER NOT NULL,
|
||||
chunks INTEGER NOT NULL,
|
||||
indexed INTEGER NOT NULL
|
||||
);
|
||||
CREATE TABLE IF NOT EXISTS vv_chunks (
|
||||
id INTEGER PRIMARY KEY,
|
||||
path TEXT NOT NULL,
|
||||
kind TEXT NOT NULL,
|
||||
section TEXT,
|
||||
heading TEXT,
|
||||
part INTEGER,
|
||||
content TEXT NOT NULL,
|
||||
vector BLOB NOT NULL,
|
||||
indexed INTEGER NOT NULL
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS idx_chunks_path ON vv_chunks(path);
|
||||
CREATE INDEX IF NOT EXISTS idx_chunks_section ON vv_chunks(section);
|
||||
CREATE INDEX IF NOT EXISTS idx_chunks_kind ON vv_chunks(kind);
|
||||
CREATE TABLE IF NOT EXISTS vv_meta (k TEXT PRIMARY KEY, v TEXT);
|
||||
`;
|
||||
|
||||
function openDb(dbPath) {
|
||||
fs.mkdirSync(path.dirname(dbPath), { recursive: true });
|
||||
const db = new DatabaseSync(dbPath);
|
||||
db.exec('PRAGMA journal_mode = WAL;');
|
||||
db.exec('PRAGMA synchronous = NORMAL;');
|
||||
db.exec(SCHEMA);
|
||||
return db;
|
||||
}
|
||||
|
||||
// Only ever index what git tracks. Configurations/, State_Files/ and data/ are gitignored, which
|
||||
// is what makes it structurally impossible for a credential to reach the index — the files that
|
||||
// hold them were never in the repo. Do not replace this with a filesystem walk.
|
||||
function trackedFiles(root) {
|
||||
return execSync('git ls-files', { cwd: root, maxBuffer: 1 << 26 })
|
||||
.toString().trim().split('\n')
|
||||
.filter(Boolean)
|
||||
.filter(f => classify(f) !== 'other');
|
||||
}
|
||||
|
||||
async function embedBatch(url, model, inputs, timeoutMs) {
|
||||
const ctl = AbortSignal.timeout(timeoutMs);
|
||||
const res = await fetch(`${url.replace(/\/$/, '')}/api/embed`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ model, input: inputs }),
|
||||
signal: ctl,
|
||||
});
|
||||
if (!res.ok) throw new Error(`embed HTTP ${res.status}: ${(await res.text()).slice(0, 200)}`);
|
||||
const j = await res.json();
|
||||
if (!j.embeddings || j.embeddings.length !== inputs.length)
|
||||
throw new Error(`embed returned ${j.embeddings ? j.embeddings.length : 0} of ${inputs.length}`);
|
||||
return j.embeddings;
|
||||
}
|
||||
|
||||
function toBlob(vec) {
|
||||
return Buffer.from(Float32Array.from(vec).buffer);
|
||||
}
|
||||
|
||||
async function buildIndex(opts) {
|
||||
const {
|
||||
root, dbPath, url, model,
|
||||
batch = 32, timeout = 120000, force = false, dryRun = false,
|
||||
onProgress = () => {},
|
||||
} = opts;
|
||||
|
||||
const db = dryRun ? null : openDb(dbPath);
|
||||
const now = Math.floor(Date.now() / 1000);
|
||||
|
||||
const known = new Map();
|
||||
if (db) for (const r of db.prepare('SELECT path, mtime FROM vv_files').all()) known.set(r.path, r.mtime);
|
||||
|
||||
const files = trackedFiles(root);
|
||||
const present = new Set(files);
|
||||
|
||||
// Files that left the repo must leave the index with them.
|
||||
let removed = 0;
|
||||
if (db && !force) {
|
||||
for (const p of known.keys()) {
|
||||
if (!present.has(p)) {
|
||||
db.prepare('DELETE FROM vv_chunks WHERE path = ?').run(p);
|
||||
db.prepare('DELETE FROM vv_files WHERE path = ?').run(p);
|
||||
removed++;
|
||||
}
|
||||
}
|
||||
}
|
||||
if (db && force) { db.exec('DELETE FROM vv_chunks; DELETE FROM vv_files;'); }
|
||||
|
||||
// ── Collect the chunks that actually need embedding ───────────────────────────────────────
|
||||
const pending = [];
|
||||
let skipped = 0, scanned = 0;
|
||||
|
||||
for (const rel of files) {
|
||||
const abs = path.join(root, rel);
|
||||
let st;
|
||||
try { st = fs.statSync(abs); } catch { continue; }
|
||||
const mtime = Math.floor(st.mtimeMs / 1000);
|
||||
scanned++;
|
||||
|
||||
if (!force && known.has(rel) && known.get(rel) === mtime) { skipped++; continue; }
|
||||
|
||||
let chunks = [];
|
||||
try { chunks = chunkFile(abs, rel); } catch (e) { continue; }
|
||||
// A file that yields no chunks still gets a vv_files row so it is not re-chunked every
|
||||
// run. Queued rather than written here, so every database change lands in the single
|
||||
// commit below — a run interrupted mid-embed must leave the index exactly as it was.
|
||||
pending.push({ rel, mtime, chunks });
|
||||
}
|
||||
|
||||
const totalChunks = pending.reduce((n, f) => n + f.chunks.length, 0);
|
||||
if (dryRun) {
|
||||
return { dryRun: true, scanned, skipped, removed, files: pending.length, chunks: totalChunks };
|
||||
}
|
||||
|
||||
// ── Embed in batches, write per file so an interrupted run leaves a consistent index ───────
|
||||
const flat = [];
|
||||
for (const f of pending) for (const c of f.chunks) flat.push({ f, c });
|
||||
|
||||
let done = 0, failed = 0;
|
||||
for (let i = 0; i < flat.length; i += batch) {
|
||||
const slice = flat.slice(i, i + batch);
|
||||
const inputs = slice.map(x => x.c.content);
|
||||
let vecs;
|
||||
try {
|
||||
vecs = await embedBatch(url, model, inputs, timeout);
|
||||
} catch (e) {
|
||||
failed += slice.length;
|
||||
onProgress({ done, total: flat.length, error: e.message });
|
||||
continue;
|
||||
}
|
||||
slice.forEach((x, k) => { x.c.__vec = vecs[k]; });
|
||||
done += slice.length;
|
||||
onProgress({ done, total: flat.length });
|
||||
}
|
||||
|
||||
const ins = db.prepare(
|
||||
'INSERT INTO vv_chunks (path,kind,section,heading,part,content,vector,indexed) VALUES (?,?,?,?,?,?,?,?)'
|
||||
);
|
||||
const insF = db.prepare('INSERT OR REPLACE INTO vv_files VALUES (?,?,?,?)');
|
||||
|
||||
db.exec('BEGIN');
|
||||
try {
|
||||
for (const f of pending) {
|
||||
// Nothing to index in this file at all — record it so it is not re-chunked next run.
|
||||
if (!f.chunks.length) {
|
||||
db.prepare('DELETE FROM vv_chunks WHERE path = ?').run(f.rel);
|
||||
insF.run(f.rel, f.mtime, 0, now);
|
||||
continue;
|
||||
}
|
||||
const embedded = f.chunks.filter(c => c.__vec);
|
||||
// A file whose chunks all failed to embed keeps its previous rows and its old mtime,
|
||||
// so the next run retries it rather than recording a half-indexed file as current.
|
||||
if (!embedded.length) continue;
|
||||
db.prepare('DELETE FROM vv_chunks WHERE path = ?').run(f.rel);
|
||||
for (const c of embedded) {
|
||||
ins.run(c.path, c.kind, c.section, c.heading, c.part, c.content, toBlob(c.__vec), now);
|
||||
}
|
||||
insF.run(f.rel, f.mtime, embedded.length, now);
|
||||
}
|
||||
db.prepare('INSERT OR REPLACE INTO vv_meta VALUES (?,?)').run('embed_model', model);
|
||||
db.prepare('INSERT OR REPLACE INTO vv_meta VALUES (?,?)').run('last_index', String(now));
|
||||
db.prepare('INSERT OR REPLACE INTO vv_meta VALUES (?,?)').run('dims', '768');
|
||||
db.exec('COMMIT');
|
||||
} catch (e) {
|
||||
db.exec('ROLLBACK');
|
||||
throw e;
|
||||
}
|
||||
|
||||
const stats = {
|
||||
scanned, skipped, removed,
|
||||
files: pending.length,
|
||||
chunks: done,
|
||||
failed,
|
||||
total: db.prepare('SELECT COUNT(*) n FROM vv_chunks').get().n,
|
||||
};
|
||||
db.close();
|
||||
return stats;
|
||||
}
|
||||
|
||||
module.exports = { buildIndex, openDb, toBlob, trackedFiles };
|
||||
@@ -0,0 +1,139 @@
|
||||
'use strict';
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
// Search — embed a question, score it against the index, return the best chunks.
|
||||
//
|
||||
// nomic-embed-text returns L2-normalised vectors, so cosine similarity is a plain dot product.
|
||||
// At this corpus size (~1.7k chunks x 768 dims) that is a couple of million multiply-adds —
|
||||
// under a millisecond, with no vector database and no index structure to maintain.
|
||||
//
|
||||
// Section routing is the payoff from the header audit. Every chunk knows whether it is a
|
||||
// PURPOSE, a DESIGN PRINCIPLES, an OPERATIONAL SAFEGUARDS and so on, so a question's shape can
|
||||
// steer retrieval before similarity is even considered. It is applied as a score boost rather
|
||||
// than a hard filter — intent detection is a heuristic, and a heuristic should not be able to
|
||||
// exclude the one chunk that actually holds the answer.
|
||||
// ═══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
const { DatabaseSync } = require('node:sqlite');
|
||||
|
||||
// Question shape → the section most likely to answer it.
|
||||
const INTENTS = [
|
||||
{ section: 'OPERATIONAL SAFEGUARDS',
|
||||
re: /\b(safe|safety|guard|protect|prevent|fail|failure|abort|refuse|lock|root|timeout|dry.?run|what stops|what happens if|race|corrupt|data.?loss)\b/i },
|
||||
{ section: 'CONFIGURATION',
|
||||
re: /\b(variable|var|setting|conf|config|threshold|toggle|which key|what controls|where is .* set|default value|env)\b/i },
|
||||
{ section: 'RUNTIME MODES',
|
||||
re: /\b(flag|argument|option|--\w+|how do i run|invoke|cli|command line|status mode|usage)\b/i },
|
||||
{ section: 'DESIGN PRINCIPLES',
|
||||
re: /\b(why|rationale|reason|design|decision|deliberate|intentional|on purpose|trade.?off|chose|approach)\b/i },
|
||||
{ section: 'OPERATIONAL MODEL',
|
||||
re: /\b(how does .* work|flow|sequence|order|tier|lifecycle|state machine|when does)\b/i },
|
||||
{ section: 'EXPORTS',
|
||||
re: /\b(function|export|api surface|what does .* provide|helper|vv_\w+)\b/i },
|
||||
{ section: 'PURPOSE',
|
||||
re: /\b(what is|what does .* do|purpose|responsible for|job of)\b/i },
|
||||
];
|
||||
|
||||
const SECTION_BOOST = 0.06; // enough to reorder near-ties, not enough to beat a real match
|
||||
const KIND_BOOST = 0.02; // docs answer "how do I" better than a script header does
|
||||
|
||||
function detectIntent(q) {
|
||||
const hits = [];
|
||||
for (const i of INTENTS) if (i.re.test(q)) hits.push(i.section);
|
||||
return hits;
|
||||
}
|
||||
|
||||
function blobToVec(buf) {
|
||||
const b = Buffer.from(buf);
|
||||
return new Float32Array(b.buffer, b.byteOffset, b.length / 4);
|
||||
}
|
||||
|
||||
function dot(a, b) {
|
||||
let s = 0;
|
||||
for (let i = 0; i < a.length; i++) s += a[i] * b[i];
|
||||
return s;
|
||||
}
|
||||
|
||||
async function embedQuery(url, model, text, timeoutMs = 60000) {
|
||||
const res = await fetch(`${url.replace(/\/$/, '')}/api/embed`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ model, input: text }),
|
||||
signal: AbortSignal.timeout(timeoutMs),
|
||||
});
|
||||
if (!res.ok) throw new Error(`embed HTTP ${res.status}`);
|
||||
const j = await res.json();
|
||||
if (!j.embeddings || !j.embeddings[0]) throw new Error('embed returned no vector');
|
||||
return Float32Array.from(j.embeddings[0]);
|
||||
}
|
||||
|
||||
// Keep at most `perFile` chunks from any one file, so a single large document cannot fill the
|
||||
// entire context window and crowd out a better answer living somewhere else.
|
||||
function diversify(rows, k, perFile) {
|
||||
const seen = new Map();
|
||||
const out = [];
|
||||
for (const r of rows) {
|
||||
const n = seen.get(r.path) || 0;
|
||||
if (n >= perFile) continue;
|
||||
seen.set(r.path, n + 1);
|
||||
out.push(r);
|
||||
if (out.length >= k) break;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
async function search(opts) {
|
||||
const {
|
||||
dbPath, url, embedModel, query,
|
||||
k = 8, perFile = 3, section = null, kind = null, minScore = 0.0,
|
||||
} = opts;
|
||||
|
||||
const db = new DatabaseSync(dbPath, { readOnly: true });
|
||||
|
||||
let sql = 'SELECT id,path,kind,section,heading,part,content,vector FROM vv_chunks';
|
||||
const where = [], args = [];
|
||||
if (section) { where.push('section = ?'); args.push(section); }
|
||||
if (kind) { where.push('kind = ?'); args.push(kind); }
|
||||
if (where.length) sql += ' WHERE ' + where.join(' AND ');
|
||||
|
||||
const rows = db.prepare(sql).all(...args);
|
||||
if (!rows.length) { db.close(); return { results: [], intents: [], scanned: 0 }; }
|
||||
|
||||
const qv = await embedQuery(url, embedModel, query);
|
||||
let intents = section ? [] : detectIntent(query);
|
||||
|
||||
// "What is Varaverk" and "what is arr_sync.sh" are not the same question, and the PURPOSE
|
||||
// intent cannot tell them apart — it fires on both and boosts every PURPOSE block in the
|
||||
// repository at once. There are a couple of hundred, each genuinely describing the purpose of
|
||||
// something, and each a short sentence containing the word Varaverk. The project's own README
|
||||
// then loses to a script that migrates storage modes, because a paragraph is more diluted
|
||||
// than a one-line summary.
|
||||
//
|
||||
// A question that names the project and no component inside it is asking about the whole, so
|
||||
// PURPOSE is precisely the wrong section to promote. Dropping only that intent, rather than
|
||||
// all of them, leaves "why was Varaverk built this way" still routed to DESIGN PRINCIPLES.
|
||||
const namesProject = /\bvaraverk\b/i.test(query);
|
||||
const namesComponent = /\b[\w.-]+\.(sh|php|js)\b|\b[A-Z][A-Z0-9]*(_[A-Z0-9]+)+\b/.test(query);
|
||||
const projectLevel = namesProject && !namesComponent;
|
||||
if (projectLevel) intents = intents.filter(s => s !== 'PURPOSE');
|
||||
|
||||
// The same question wants the top-level prose, which is what the doc kinds are.
|
||||
const wantDoc = projectLevel
|
||||
|| /\b(how do i|steps|procedure|setup|install|troubleshoot|guide)\b/i.test(query);
|
||||
|
||||
const scored = rows.map(r => {
|
||||
let s = dot(qv, blobToVec(r.vector));
|
||||
if (intents.includes(r.section)) s += SECTION_BOOST;
|
||||
if (wantDoc && (r.kind === 'manual' || r.kind === 'readme')) s += KIND_BOOST;
|
||||
return {
|
||||
id: r.id, path: r.path, kind: r.kind, section: r.section,
|
||||
heading: r.heading, part: r.part, content: r.content, score: s,
|
||||
};
|
||||
});
|
||||
|
||||
scored.sort((a, b) => b.score - a.score);
|
||||
const kept = diversify(scored.filter(r => r.score >= minScore), k, perFile);
|
||||
db.close();
|
||||
return { results: kept, intents, scanned: rows.length };
|
||||
}
|
||||
|
||||
module.exports = { search, detectIntent, blobToVec, dot, embedQuery, INTENTS };
|
||||
@@ -0,0 +1,670 @@
|
||||
# ━━━━━ ARRS STACK — Manual ━━━━━
|
||||
|
||||
Config reference, procedures, operational workflows.
|
||||
For overview see README-Arrs_Stack.md. For per-script detail see script headers.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ ARR CLEANUP — FILE CLASSIFICATION ━━━
|
||||
|
||||
Every file found on disk during an arr cleanup run falls into exactly one category:
|
||||
|
||||
```
|
||||
TRACKED → arr API returned this exact path → leave it alone
|
||||
PROTECTED → matches ARR_PROTECTED_PATTERNS → never delete
|
||||
ORPHAN → media extension, not tracked, old enough → delete
|
||||
JUNK → not a media extension, not protected → delete (any age)
|
||||
RECENT → not tracked, under ARR_ORPHAN_AGE days → skip (may be mid-import)
|
||||
```
|
||||
|
||||
**Why protected patterns are critical:** arrs generate artwork (`*.jpg`), metadata
|
||||
(`*.nfo`), and subtitles/lyrics that do NOT appear in the tracked file API response.
|
||||
Without protection, these would be classified as orphans and deleted — removing cover art
|
||||
from every album, every movie poster, every TV show thumbnail. Requires a full rescan
|
||||
to recover. Never remove artwork extensions from protected patterns.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ ARR CLEANUP — SAFETY LAYERS ━━━
|
||||
|
||||
All 7 layers must pass before any file is touched. There is no way to push through a
|
||||
failed safety check without the explicit override flag.
|
||||
|
||||
```
|
||||
1. Container running + healthy — a stopped container has an empty API
|
||||
2. API reachable — no API = no tracked file list = everything looks orphaned
|
||||
3. API version matches — major version must match tested version in master.conf
|
||||
4. Item count > 0 — no artists/series/movies = something is wrong with DB
|
||||
5. Tracked file count > 0 — empty response = everything would be deleted
|
||||
6. Tracked count >= MIN_TRACKED_PCT — dramatic drop from last run = abort and alert
|
||||
7. Deletion size < MAX_DELETE_GB — last line of defense against misconfigured root path
|
||||
```
|
||||
|
||||
Layer 7 is the catastrophic failure prevention. A misconfigured root path — pointing
|
||||
cleanup at the wrong directory — means the API returns zero tracked files for a root
|
||||
that actually contains thousands. Everything walks as an orphan. Everything gets deleted.
|
||||
`LIDARR/SONARR/RADARR_MAX_DELETE_GB` requires `--i-know-what-im-doing` to proceed past it.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ CONFIGURATION — master.conf ━━━
|
||||
|
||||
### Lidarr Cleanup Thresholds
|
||||
|
||||
```bash
|
||||
LIDARR_ORPHAN_AGE=7 # days — files newer than this are RECENT (mid-import window)
|
||||
LIDARR_MIN_TRACKED_PCT=80 # abort if API returns < 80% of last known count
|
||||
LIDARR_MAX_DELETE_GB=5 # require --i-know-what-im-doing above this
|
||||
LIDARR_IMPORT_SCAN_TIMEOUT=600 # seconds to wait for pre-flight import scan
|
||||
LIDARR_VERSION_MAJOR=3 # expected Lidarr major version (API safety check)
|
||||
LIDARR_LOCK_WARN_AGE=3600 # 1hr — large libraries take time, not stuck
|
||||
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
||||
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
||||
LIDARR_TRACKED_COUNT_FILE="$DATA_DIR/lidarr_tracked.count" # persistent baseline
|
||||
ARR_CLEANUP_STATS="$DATA_DIR/arr_cleanup_stats.db" # read by coffee report
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Lidarr Release Fixer
|
||||
|
||||
```bash
|
||||
LIDARR_RELEASE_FIXER_ENABLED=true # set false to disable without removing from job list
|
||||
```
|
||||
|
||||
No additional thresholds — uses existing `LIDARR_URL`, `LIDARR_API_KEY`,
|
||||
`LIDARR_MUSIC_ROOT`, `LIDARR_PATH_MAP`, and `LIDARR_VERSION_MAJOR` from host*.conf and
|
||||
master.conf. Reads FLAC (vorbis comment block type 4) and MP3 (ID3v2 TXXX frame) tags.
|
||||
|
||||
---
|
||||
|
||||
### Sonarr Cleanup Thresholds
|
||||
|
||||
```bash
|
||||
SONARR_ORPHAN_AGE=7
|
||||
SONARR_MAX_DELETE_GB=50
|
||||
SONARR_IMPORT_SCAN_TIMEOUT=600
|
||||
SONARR_VERSION_MAJOR=4
|
||||
SONARR_EXTENSIONS=("mkv" "mp4" "avi" "m4v" "ts" "wmv" "mov")
|
||||
SONARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.ass" "*.ssa")
|
||||
```
|
||||
|
||||
Note: `*.ts` IS in extensions — transport stream is used for Live TV recordings tracked
|
||||
by Sonarr. Orphaned `.ts` recordings should be cleaned like any other orphaned episode.
|
||||
|
||||
---
|
||||
|
||||
### Radarr Cleanup Thresholds
|
||||
|
||||
```bash
|
||||
RADARR_ORPHAN_AGE=7
|
||||
RADARR_MAX_DELETE_GB=50
|
||||
RADARR_IMPORT_SCAN_TIMEOUT=600
|
||||
RADARR_VERSION_MAJOR=6
|
||||
RADARR_EXTENSIONS=("mkv" "mp4" "avi" "m4v" "wmv" "mov")
|
||||
RADARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.ass" "*.ssa")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Arr Sync
|
||||
|
||||
```bash
|
||||
ARR_SYNC_ENABLED=true
|
||||
ARR_SYNC_BLOCKLIST="$DATA_DIR/arr_sync_blocklist.tsv" # tombstone file
|
||||
ARR_SYNC_CONNECT_TIMEOUT=10 # SSH connect timeout in seconds
|
||||
ARR_SYNC_API_TIMEOUT=60 # curl API call timeout in seconds
|
||||
DOCKER_APPDATA_BASE=/mnt/user/appdata
|
||||
ARR_SYNC_LIDARR_PORT=8686
|
||||
ARR_SYNC_SONARR_PORT=8989
|
||||
ARR_SYNC_RADARR_PORT=7878
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Arr Recovery
|
||||
|
||||
```bash
|
||||
ARR_IMPORT_RECOVERY_AGE=6 # hours — items newer than this are skipped
|
||||
SONARR_VERSION_MAJOR=4
|
||||
RADARR_VERSION_MAJOR=6
|
||||
LIDARR_VERSION_MAJOR=3
|
||||
ARR_RECOVERY_STATS="$DATA_DIR/arr_recovery_stats.db" # read by coffee report
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### TMDb / TVDB Removed
|
||||
|
||||
```bash
|
||||
RADARR_DROPPED_ADD_EXCLUSION=true # add removed movies to Radarr import exclusion
|
||||
SONARR_DROPPED_ADD_EXCLUSION=true # add removed series to Sonarr import exclusion
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Lidarr Missing Art
|
||||
|
||||
```bash
|
||||
FANART_API_KEY="your-fanart-tv-api-key"
|
||||
LASTFM_API_KEY="your-lastfm-api-key"
|
||||
LIDARR_ART_MIN_SIZE=5000 # minimum valid download size in bytes
|
||||
LIDARR_ART_MAX_PARALLEL=4 # concurrent background download jobs
|
||||
LIDARR_ART_RETRIES=2 # download retry attempts per image
|
||||
LIDARR_ART_SLEEP_BETWEEN=1 # seconds between fanart.tv API calls (rate limit)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Lidarr Discovery
|
||||
|
||||
```bash
|
||||
LIDARR_DISCOVERY_THRESHOLD=70 # minimum score for Stage 1 seeds and Stage 2 adds
|
||||
LIDARR_DISCOVERY_LOOKBACK_DAYS=7 # Emby play history window in days
|
||||
LIDARR_DISCOVERY_MIN_PLAYS=3 # min plays before an artist is evaluated as a seed
|
||||
LIDARR_DISCOVERY_MAX_ADDS=5 # max seeds (Stage 1) and max adds (Stage 2) per run
|
||||
LIDARR_DISCOVERY_USER_CAP_PCT=35 # max % any one user contributes to play weight
|
||||
LIDARR_DISCOVERY_REJECT_COOLDOWN=30 # days before re-evaluating a Stage 2 reject
|
||||
LIDARR_DISCOVERY_HISTORY="$DATA_DIR/lidarr_discovery_history.db"
|
||||
```
|
||||
|
||||
Requires `HOST*_LASTFM_API_KEY` in `host*.conf`.
|
||||
|
||||
---
|
||||
|
||||
### Radarr Discovery
|
||||
|
||||
```bash
|
||||
RADARR_DISCOVERY_THRESHOLD=52 # minimum score to add a candidate
|
||||
RADARR_DISCOVERY_LOOKBACK_DAYS=30 # Emby watch history window in days
|
||||
RADARR_DISCOVERY_MAX_SEEDS=5 # max seed movies from Stage 1
|
||||
RADARR_DISCOVERY_MAX_ADDS=5 # max movies to add per run
|
||||
RADARR_DISCOVERY_MIN_VOTE_COUNT=100 # min TMDB votes for a candidate
|
||||
RADARR_DISCOVERY_MIN_RATING=60 # min TMDB vote_average × 10 (60 = 6.0/10)
|
||||
RADARR_DISCOVERY_REJECT_COOLDOWN=60 # days before re-evaluating a rejected movie
|
||||
RADARR_DISCOVERY_SEED_LIBRARIES=("Movies")
|
||||
RADARR_DISCOVERY_HISTORY="$DATA_DIR/radarr_discovery_history.db"
|
||||
```
|
||||
|
||||
Requires `HOST*_TMDB_API_KEY` in `host*.conf`.
|
||||
|
||||
---
|
||||
|
||||
### Sonarr Discovery
|
||||
|
||||
```bash
|
||||
SONARR_DISCOVERY_THRESHOLD=52 # minimum score to add a candidate
|
||||
SONARR_DISCOVERY_LOOKBACK_DAYS=14 # Emby episode history window in days
|
||||
SONARR_DISCOVERY_MAX_SEEDS=5 # max seed series from Stage 1
|
||||
SONARR_DISCOVERY_MAX_ADDS=3 # max shows to add per run (TV is a larger commitment)
|
||||
SONARR_DISCOVERY_MIN_VOTE_COUNT=50 # min TMDB votes for a candidate
|
||||
SONARR_DISCOVERY_MIN_RATING=65 # min TMDB vote_average × 10 (65 = 6.5/10)
|
||||
SONARR_DISCOVERY_REJECT_COOLDOWN=60 # days before re-evaluating a rejected show
|
||||
SONARR_DISCOVERY_USER_EPISODE_CAP=8 # max episodes per user in seed scoring
|
||||
SONARR_DISCOVERY_MONITOR_MODE="all" # "all" = all seasons monitored; "future" = upcoming only
|
||||
SONARR_DISCOVERY_HISTORY="$DATA_DIR/sonarr_discovery_history.db"
|
||||
```
|
||||
|
||||
Requires `HOST*_TMDB_API_KEY` in `host*.conf`.
|
||||
|
||||
> **MONITOR_MODE note:** Use `"all"` (default) to have Sonarr search all existing seasons
|
||||
> after adding a show. `"future"` only marks upcoming seasons as monitored — shows where
|
||||
> all seasons have already aired will appear unmonitored and Sonarr will not search for them.
|
||||
|
||||
---
|
||||
|
||||
### Upgrade Webhook
|
||||
|
||||
```bash
|
||||
WEBHOOK_PORT=7821 # 0 = disable listener
|
||||
WEBHOOK_SECRET="" # auto-generated on first start if empty
|
||||
```
|
||||
|
||||
When `WEBHOOK_PORT=0`, `start_webhook_listener.sh` exits cleanly and no listener starts.
|
||||
When `WEBHOOK_SECRET` is empty, a 32-byte hex secret is generated on first start and
|
||||
written back to `master.conf`. Run `Tools/webhook_setup.sh` to register the URL in arrs.
|
||||
|
||||
Webhook log: `/var/log/varaverk/upgrade_webhook.log`
|
||||
|
||||
---
|
||||
|
||||
## ━━━ CONFIGURATION — host*.conf ━━━
|
||||
|
||||
```bash
|
||||
# Arr connection details — must match arr settings exactly
|
||||
HOST1_LIDARR_URL="http://192.168.50.2:8686"
|
||||
HOST1_LIDARR_API_KEY="..."
|
||||
HOST1_LIDARR_MUSIC_ROOT="/mnt/user/Music"
|
||||
HOST1_LIDARR_PATH_MAP="" # container→host path translation if needed
|
||||
|
||||
HOST1_SONARR_URL="http://192.168.50.2:8989"
|
||||
HOST1_SONARR_API_KEY="..."
|
||||
HOST1_SONARR_TV_ROOT="/mnt/user/Tv_Shows"
|
||||
HOST1_SONARR_PATH_MAP=""
|
||||
|
||||
HOST1_RADARR_URL="http://192.168.50.2:7878"
|
||||
HOST1_RADARR_API_KEY="..."
|
||||
HOST1_RADARR_MOVIES_ROOT="/mnt/user/Movies"
|
||||
HOST1_RADARR_PATH_MAP=""
|
||||
|
||||
HOST1_EMBY_URL="http://192.168.50.2:8096"
|
||||
HOST1_EMBY_API_KEY="..."
|
||||
```
|
||||
|
||||
> **LIDARR/SONARR/RADARR_MUSIC/TV/MOVIES_ROOT must exactly match the Root Folder path in
|
||||
> the arr's own settings.** Arr UI → Settings → Media Management → Root Folders.
|
||||
> A mismatch means every file on disk looks untracked — all appear as orphans.
|
||||
> MAX_DELETE_GB is the only thing standing between a path mismatch and losing your library.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SAFE TESTING PROCEDURE ━━━
|
||||
|
||||
> **The arr cleanup scripts permanently delete files.** There is no recycle bin, no undo.
|
||||
> Follow this procedure on first use, after any root path change, after any API key change,
|
||||
> and after any significant arr library change.
|
||||
|
||||
### Step 1 — Dry Run With Full Logging
|
||||
|
||||
```bash
|
||||
lidarr_release_fixer.sh --dry-run --log
|
||||
lidarr_cleanup.sh --dry-run --log
|
||||
sonarr_cleanup.sh --dry-run --log
|
||||
radarr_cleanup.sh --dry-run --log
|
||||
```
|
||||
|
||||
### Step 2 — Review the Output
|
||||
|
||||
```
|
||||
Release fixer:
|
||||
Are the "would fix" albums expected?
|
||||
→ Wrong release UUIDs switching to correct ones is expected behaviour
|
||||
→ If everything is already correct, nothing to do
|
||||
|
||||
Cleanup:
|
||||
Are TRACKED files the ones you expect?
|
||||
→ Known arr-managed files should show as TRACKED
|
||||
→ If they show as ORPHAN, the root path is wrong — STOP
|
||||
|
||||
Is the ORPHAN count reasonable?
|
||||
→ Healthy cleanup removes dozens to hundreds, not tens of thousands
|
||||
→ Large count = stop, investigate root path before proceeding
|
||||
|
||||
Are PROTECTED patterns working?
|
||||
→ Artwork (*.jpg) and subtitles (*.srt) must show as PROTECTED
|
||||
→ If they show as ORPHAN, check PROTECTED_PATTERNS config
|
||||
|
||||
Are RECENT files being correctly skipped?
|
||||
→ Files downloaded in the last 7 days should show as RECENT, not ORPHAN
|
||||
```
|
||||
|
||||
### Step 3 — Check Numbers if Something Looks Wrong
|
||||
|
||||
```bash
|
||||
# Root path mismatch? Compare these:
|
||||
# Lidarr UI: Settings → Media Management → Root Folders
|
||||
# Sonarr UI: Settings → Media Management → Root Folders
|
||||
# Radarr UI: Settings → Media Management → Root Folders
|
||||
# Must exactly match LIDARR_MUSIC_ROOT / SONARR_TV_ROOT / RADARR_MOVIES_ROOT
|
||||
|
||||
# Is the arr running?
|
||||
docker ps | grep -E "Lidarr|Sonarr|Radarr"
|
||||
```
|
||||
|
||||
### Step 4 — Run Live
|
||||
|
||||
```bash
|
||||
# Only after dry run review passes.
|
||||
lidarr_release_fixer.sh
|
||||
lidarr_cleanup.sh
|
||||
sonarr_cleanup.sh
|
||||
radarr_cleanup.sh
|
||||
```
|
||||
|
||||
### Step 5 — Verify in Arr UI
|
||||
|
||||
```
|
||||
Library count — should not have dropped significantly
|
||||
Missing files — check if any monitored content shows as missing
|
||||
Emby library — should show no ghost entries (notify_emby_scan handles this automatically)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ PROCEDURES ━━━
|
||||
|
||||
### Adding a New Arr
|
||||
|
||||
```bash
|
||||
# 1. Copy radarr_cleanup.sh as template
|
||||
cp Arrs_Stack/radarr_cleanup.sh Arrs_Stack/readarr_cleanup.sh
|
||||
|
||||
# 2. Replace RADARR_ prefix with READARR_ throughout
|
||||
# Update API endpoint, tracked file API path, extension list, protected patterns
|
||||
|
||||
# 3. Add to host*.conf
|
||||
HOST1_READARR_URL="http://192.168.50.2:8787"
|
||||
HOST1_READARR_API_KEY="your-api-key"
|
||||
HOST1_READARR_BOOKS_ROOT="/mnt/user/Books"
|
||||
|
||||
# 4. Add thresholds to master.conf
|
||||
READARR_ORPHAN_AGE=7
|
||||
READARR_MAX_DELETE_GB=50
|
||||
READARR_EXTENSIONS=("epub" "pdf" "mobi" "azw3" "cbz" "cbr")
|
||||
READARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo")
|
||||
|
||||
# 5. Add to DAILY_MAINTENANCE_SCRIPTS in master.conf
|
||||
"Arrs_Stack/readarr_cleanup.sh"
|
||||
```
|
||||
|
||||
Run `--dry-run --log` before scheduling.
|
||||
|
||||
---
|
||||
|
||||
### Managing the Arr Sync Blocklist
|
||||
|
||||
```bash
|
||||
# Add item to blocklist (removes from all arrs + tombstones the ID)
|
||||
arr_sync.sh --blocklist-add lidarr <musicbrainz-artist-id> "reason"
|
||||
arr_sync.sh --blocklist-add sonarr <tvdb-series-id> "reason"
|
||||
arr_sync.sh --blocklist-add radarr <tmdb-movie-id> "reason"
|
||||
|
||||
# Remove from blocklist (un-tombstones the ID — does NOT re-add to arrs)
|
||||
arr_sync.sh --blocklist-remove lidarr <id>
|
||||
|
||||
# View all blocklisted IDs
|
||||
arr_sync.sh --blocklist-list
|
||||
```
|
||||
|
||||
`--blocklist-add` is the only destructive operation — it simultaneously:
|
||||
1. Writes the tombstone entry to the blocklist TSV file
|
||||
2. Deletes the item from the local arr API (no file deletion)
|
||||
3. SSHes each remote node and deletes from their arr API
|
||||
|
||||
Files become orphans on all nodes — arr_cleanup removes them on the next run.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ TROUBLESHOOTING ━━━
|
||||
|
||||
### Arr Cleanup Deleting Files It Shouldn't
|
||||
|
||||
```
|
||||
1. Check the protected patterns — artwork and subtitles must be listed
|
||||
LIDARR_PROTECTED_PATTERNS / SONARR_PROTECTED_PATTERNS / RADARR_PROTECTED_PATTERNS
|
||||
|
||||
2. Check the root path matches arr settings exactly
|
||||
Run: lidarr_cleanup.sh --status (shows configured root path)
|
||||
Compare: Lidarr UI → Settings → Media Management → Root Folders
|
||||
|
||||
3. Check if files are truly orphaned
|
||||
Run: lidarr_cleanup.sh --dry-run --log
|
||||
Look for the specific file — verify it shows ORPHAN, not PROTECTED or TRACKED
|
||||
```
|
||||
|
||||
### Lidarr Files Present But Not Importing
|
||||
|
||||
```
|
||||
Symptom: album folder exists with properly tagged files, Lidarr shows 0 tracks imported.
|
||||
Cause: Lidarr selected the wrong MusicBrainz release edition. The file's MBID doesn't
|
||||
match Lidarr's foreignReleaseId, so track ID lookup fails.
|
||||
|
||||
Fix:
|
||||
lidarr_release_fixer.sh --dry-run --log
|
||||
→ Shows which albums would be corrected and which release UUIDs would change
|
||||
|
||||
lidarr_release_fixer.sh
|
||||
→ Applies corrections and queues RefreshArtist for affected artists
|
||||
|
||||
If the file MBID doesn't match any Lidarr release (shows "MBID not in list"):
|
||||
→ File was tagged from a source Lidarr doesn't know about (unofficial, re-tagged)
|
||||
→ Manual import or re-tag with correct MBID
|
||||
```
|
||||
|
||||
### Arr Cleanup Aborting at Safety Layer 6 (Tracked Count Drop)
|
||||
|
||||
```
|
||||
API returned far fewer tracked files than last run.
|
||||
Possible causes:
|
||||
- Arr database was recently rebuilt from scratch
|
||||
- Large manual library removal
|
||||
- Path map mismatch after arr migration
|
||||
|
||||
If intentional (library intentionally reduced):
|
||||
Delete LIDARR_TRACKED_COUNT_FILE to reset the baseline
|
||||
Run cleanup once — it will establish a new baseline
|
||||
|
||||
If unintentional:
|
||||
Investigate before proceeding — the arr may have a problem
|
||||
```
|
||||
|
||||
### Arr Sync Not Picking Up New Content
|
||||
|
||||
```
|
||||
Is ARR_SYNC_ENABLED=true in master.conf?
|
||||
|
||||
Can this host SSH to the remote without password?
|
||||
→ ssh -i [SSH_KEY] root@[remote-tailscale-ip] "hostname"
|
||||
|
||||
Is the arr accessible on the remote?
|
||||
→ arr_sync.sh --status (shows each node's arr reachability)
|
||||
→ arr_sync.sh --log (verbose output per-node, per-arr)
|
||||
|
||||
Is the item in the blocklist?
|
||||
→ arr_sync.sh --blocklist-list
|
||||
```
|
||||
|
||||
### Emby Still Showing Ghost Entries After Cleanup
|
||||
|
||||
```
|
||||
notify_emby_scan() is called automatically after every arr cleanup deletion.
|
||||
If ghosts persist:
|
||||
1. Is Emby's API responding?
|
||||
curl -s "http://[emby-ip]:8096/System/Info/Public"
|
||||
|
||||
2. Is EMBY_URL / EMBY_API_KEY correct in host*.conf?
|
||||
Run: sonarr_cleanup.sh --status (shows Emby config)
|
||||
|
||||
3. Trigger manually in Emby:
|
||||
Library → Manage Library → Clean Missing Files
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ OUTPUT TIERS ━━━
|
||||
|
||||
All scripts have two output levels controlled by `--log`.
|
||||
|
||||
Without `--log`, each script processes silently and always concludes with a summary
|
||||
block: identity, duration, counts (files removed, items added, arrs cleaned), and a
|
||||
status line. Warnings and errors are always visible.
|
||||
|
||||
With `--log`, per-item detail appears: individual titles being processed, API query
|
||||
progress, per-node sync results, and per-file examination output. Use this when
|
||||
debugging unexpected results or validating configuration before the first scheduled run.
|
||||
|
||||
Dry-run output follows the same tiers — `--dry-run` alone shows the summary of what
|
||||
would happen; `--dry-run --log` shows the full per-item preview list.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ FLAG REFERENCE ━━━
|
||||
|
||||
### lidarr_release_fixer.sh
|
||||
|
||||
`lidarr_release_fixer.sh`
|
||||
Read MUSICBRAINZ_ALBUMID from FLAC/MP3 files, correct Lidarr release selection, queue RefreshArtist.
|
||||
|
||||
`lidarr_release_fixer.sh --dry-run`
|
||||
Show which albums would be corrected and the before/after release UUIDs. No API writes.
|
||||
|
||||
`lidarr_release_fixer.sh --status`
|
||||
Show configuration, Lidarr URL, enabled state.
|
||||
|
||||
`lidarr_release_fixer.sh --log`
|
||||
Verbose per-album output — shows every album checked, not just those corrected.
|
||||
|
||||
---
|
||||
|
||||
### lidarr_cleanup.sh / sonarr_cleanup.sh / radarr_cleanup.sh
|
||||
|
||||
`[script] --dry-run --log`
|
||||
Preview every classification decision. **Always run this first.** See Safe Testing Procedure.
|
||||
|
||||
`[script]`
|
||||
Live run — deletes confirmed orphans and junk, triggers Emby clean.
|
||||
|
||||
`[script] --log`
|
||||
Live run with verbose per-file output.
|
||||
|
||||
`[script] --status`
|
||||
Show configuration, API status, tracked file count, and last run stats.
|
||||
|
||||
`[script] --i-know-what-im-doing`
|
||||
Bypass the MAX_DELETE_GB size threshold. Required when deletion exceeds the configured
|
||||
limit. Long flag name is intentional — cannot be added accidentally.
|
||||
|
||||
`[script] --skip-age-check`
|
||||
Bypass the ORPHAN_AGE age check. Deletes RECENT files too — files that are under the
|
||||
age threshold. Use when you know recent downloads are actually orphans.
|
||||
|
||||
`[script] --i-know-what-im-doing --skip-age-check`
|
||||
**NUCLEAR MODE** — age check and size threshold both bypassed. Deletes on first pass.
|
||||
Use when you want a clean one-pass wipe of everything the arr doesn't track.
|
||||
No recovery possible after deletion.
|
||||
|
||||
---
|
||||
|
||||
### arrs_failed_stalled_recovery.sh
|
||||
|
||||
`arrs_failed_stalled_recovery.sh`
|
||||
Check all configured arrs for failed imports and stalled downloads. Blocklist + remove +
|
||||
re-search for each problem item.
|
||||
|
||||
`arrs_failed_stalled_recovery.sh --dry-run`
|
||||
Show what would be actioned per arr without making any changes.
|
||||
|
||||
`arrs_failed_stalled_recovery.sh --status`
|
||||
Show configuration, arr reachability, and last recovery stats.
|
||||
|
||||
`arrs_failed_stalled_recovery.sh --log`
|
||||
Verbose output per item per arr.
|
||||
|
||||
---
|
||||
|
||||
### arr_sync.sh
|
||||
|
||||
`arr_sync.sh`
|
||||
Sync all arr types across all configured nodes.
|
||||
|
||||
`arr_sync.sh --dry-run`
|
||||
Show what would be added/removed on each node without making changes.
|
||||
|
||||
`arr_sync.sh --status`
|
||||
Show node configuration, arr reachability, and blocklist count.
|
||||
|
||||
`arr_sync.sh --log`
|
||||
Verbose per-node, per-arr output.
|
||||
|
||||
`arr_sync.sh --blocklist-add [arr] [id] "[reason]"`
|
||||
Remove item from all arrs and tombstone the ID. See Procedures above.
|
||||
|
||||
`arr_sync.sh --blocklist-remove [arr] [id]`
|
||||
Remove tombstone — does NOT re-add item to arrs.
|
||||
|
||||
`arr_sync.sh --blocklist-list`
|
||||
Show all tombstoned IDs.
|
||||
|
||||
---
|
||||
|
||||
### radarr_tmdb_removed.sh / sonarr_tvdb_removed.sh
|
||||
|
||||
`[script]`
|
||||
Remove records for entries with status="deleted" (dropped from upstream database).
|
||||
Files are kept. Import exclusion is added.
|
||||
|
||||
`[script] --delete-files`
|
||||
Also delete associated files from disk. Most dropped entries have no files — they were
|
||||
announced movies/series that were never downloaded.
|
||||
|
||||
`[script] --dry-run`
|
||||
Preview what would be removed without making changes.
|
||||
|
||||
`[script] --status`
|
||||
Show arr connection status and current count of dropped entries.
|
||||
|
||||
`[script] --log`
|
||||
Verbose per-entry output.
|
||||
|
||||
---
|
||||
|
||||
### lidarr_missing_art.sh
|
||||
|
||||
`lidarr_missing_art.sh`
|
||||
Fetch all missing album and artist artwork from fanart.tv and fallback sources.
|
||||
Never overwrites existing files.
|
||||
|
||||
`lidarr_missing_art.sh --dry-run`
|
||||
Show what would be downloaded without writing any files.
|
||||
|
||||
`lidarr_missing_art.sh --status`
|
||||
Show configuration and API key status.
|
||||
|
||||
`lidarr_missing_art.sh --log`
|
||||
Verbose per-album, per-artist output.
|
||||
|
||||
---
|
||||
|
||||
### playback_aware_lidarr_discovery.sh
|
||||
|
||||
`playback_aware_lidarr_discovery.sh`
|
||||
Score Emby play history, run Last.fm getSimilar on top artists, add candidates above
|
||||
threshold to Lidarr. Triggers ArtistSearch immediately after each successful add.
|
||||
|
||||
`playback_aware_lidarr_discovery.sh --dry-run`
|
||||
Score and rank all Stage 1 seeds and Stage 2 candidates. No Lidarr API calls. No writes
|
||||
to history file. Shows exactly what would be added and at what score.
|
||||
|
||||
`playback_aware_lidarr_discovery.sh --status`
|
||||
Show config values, history file path and size, and API key status.
|
||||
|
||||
`playback_aware_lidarr_discovery.sh --log`
|
||||
Verbose per-artist scoring output for both stages.
|
||||
|
||||
---
|
||||
|
||||
### playback_aware_radarr_discovery.sh
|
||||
|
||||
`playback_aware_radarr_discovery.sh`
|
||||
Score recently watched Emby movies, run TMDB recommendations on seeds, add candidates
|
||||
above threshold to Radarr. Triggers MoviesSearch immediately after each successful add.
|
||||
|
||||
`playback_aware_radarr_discovery.sh --dry-run`
|
||||
Score and rank all Stage 1 seeds and Stage 2 candidates. No Radarr API calls. No writes
|
||||
to history file.
|
||||
|
||||
`playback_aware_radarr_discovery.sh --status`
|
||||
Show config values, history file path and size, and API key status.
|
||||
|
||||
`playback_aware_radarr_discovery.sh --log`
|
||||
Verbose per-movie scoring output for both stages.
|
||||
|
||||
---
|
||||
|
||||
### playback_aware_sonarr_discovery.sh
|
||||
|
||||
`playback_aware_sonarr_discovery.sh`
|
||||
Score recently watched Emby series (weighted by user diversity), run TMDB TV
|
||||
recommendations on seeds, add candidates above threshold to Sonarr. Triggers SeriesSearch
|
||||
immediately after each successful add.
|
||||
|
||||
`playback_aware_sonarr_discovery.sh --dry-run`
|
||||
Score and rank all Stage 1 seeds and Stage 2 candidates. No Sonarr API calls. No writes
|
||||
to history file.
|
||||
|
||||
`playback_aware_sonarr_discovery.sh --status`
|
||||
Show config values, history file path and size, and API key status.
|
||||
|
||||
`playback_aware_sonarr_discovery.sh --log`
|
||||
Verbose per-series scoring output — shows user diversity, recency, and volume scores per
|
||||
seed; breadth, rating, and votes scores per candidate.
|
||||
@@ -0,0 +1,319 @@
|
||||
# ━━━━━ ARRS STACK ━━━━━
|
||||
|
||||
Lifecycle management for the arr suite (Lidarr, Sonarr, Radarr) across a multi-server
|
||||
ecosystem. Library sync so every node tracks the same content. Orphan cleanup against live
|
||||
arr APIs so deleted content actually leaves disk. Release correction so wrong MusicBrainz
|
||||
editions don't silently block Lidarr imports. Emby notified automatically after every
|
||||
deletion. Failed downloads recovered overnight. Artwork fetched continuously. Database
|
||||
hygiene for dropped upstream entries. Quality upgrades propagate to all nodes immediately
|
||||
via webhook. Weekly discovery adds new content based on what you actually play and watch.
|
||||
|
||||
> **These scripts permanently delete files and modify arr databases.** Orphan cleanup is
|
||||
> protected by multiple safety layers that must all pass before anything is touched — but
|
||||
> dry runs and log review are still the right first step on any new system or after any
|
||||
> configuration change. See Manual-Arrs_Stack.md for the safe testing procedure.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE PROBLEMS THAT BUILT THIS ━━━
|
||||
|
||||
**Deleted Shows and Removed Albums Still on Disk**
|
||||
When you remove a series from Sonarr and the delete command fails — permission issue,
|
||||
container wasn't running, path mismatch — the files stay permanently. Over years on an
|
||||
active library this accumulates significantly.
|
||||
The fix: arr cleanup scripts query the live API for every tracked file path, walk the
|
||||
disk, and delete anything absent from the API response that's old enough to be past the
|
||||
import window.
|
||||
|
||||
**Wrong MusicBrainz Release Edition Silently Blocking Lidarr Imports**
|
||||
Lidarr selects one specific release edition per album using a MusicBrainz release ID.
|
||||
When it picks the wrong edition (Brazil CD instead of US CD, Japan Digital instead of
|
||||
standard), the track IDs don't match what's embedded in the files. RescanFolders reports
|
||||
"Importing 0 tracks" even with perfectly tagged, complete files present. No error — just
|
||||
silence.
|
||||
The fix: `lidarr_release_fixer.sh` reads the MUSICBRAINZ_ALBUMID tag from each file,
|
||||
finds the matching release in Lidarr's known releases, switches the selection, and queues
|
||||
a RefreshArtist. Runs before cleanup so corrected albums are imported before the orphan
|
||||
scan ever sees them.
|
||||
|
||||
**Emby Showing Ghost Entries After Cleanup**
|
||||
After arr cleanup deletes files, Emby still shows them until its next scheduled scan —
|
||||
potentially hours later. Users see broken entries that produce "file not found" errors.
|
||||
The fix: `notify_emby_scan()` is called automatically after every deletion. Triggers
|
||||
Emby's "Clean Missing Files" task immediately.
|
||||
|
||||
**No Safety Net on Deletion Size**
|
||||
A misconfigured root path — pointing cleanup at the wrong directory — means the API
|
||||
returns zero tracked files for a root that actually contains thousands. Every file walks
|
||||
as an orphan. Everything gets deleted. This is the catastrophic failure mode.
|
||||
The fix: `LIDARR/SONARR/RADARR_MAX_DELETE_GB` — if total deletion size exceeds the
|
||||
limit, the script stops and requires `--i-know-what-im-doing` to proceed. The flag name
|
||||
is long and annoying by design. It cannot be added by accident.
|
||||
|
||||
**Failed Downloads Accumulating Silently**
|
||||
Import failures and stalled downloads sit in arr queues indefinitely. Without
|
||||
intervention they occupy queue slots, block new searches, and the item never gets
|
||||
downloaded. Checking queues manually across three arrs is tedious.
|
||||
The fix: `arrs_failed_stalled_recovery.sh` inspects all arr queues, blocklists the bad
|
||||
release, removes it, and triggers a re-search. Runs daily.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHAT THIS FOLDER DOES ━━━
|
||||
|
||||
**Library Sync**
|
||||
`arr_sync.sh` — full-mesh arr library sync across all nodes. Every node syncs with every
|
||||
other — union model, no hierarchy. Once arrs agree on what to track, rsync spreads the
|
||||
actual files.
|
||||
|
||||
**Release Correction**
|
||||
`lidarr_release_fixer.sh` — reads MUSICBRAINZ_ALBUMID from FLAC and MP3 files, matches
|
||||
against Lidarr's known releases per album, switches monitored=true to the correct
|
||||
edition, queues RefreshArtist. Runs before lidarr_cleanup.sh daily. Handles both FLAC
|
||||
(vorbis comment block) and MP3 (ID3v2 TXXX frame).
|
||||
|
||||
**Orphan Cleanup**
|
||||
`lidarr_cleanup.sh`, `sonarr_cleanup.sh`, `radarr_cleanup.sh` — API-verified orphan
|
||||
removal. Five classification categories (TRACKED/PROTECTED/ORPHAN/JUNK/RECENT), seven
|
||||
safety layers, automatic Emby notification after deletion.
|
||||
|
||||
**Database Hygiene**
|
||||
`radarr_tmdb_removed.sh`, `sonarr_tvdb_removed.sh` — remove entries that upstream
|
||||
databases have dropped (TMDb/TVDB status="deleted"). These generate health warnings in
|
||||
arrs and can never be monitored or downloaded. Most are announced-but-never-released
|
||||
entries. Files are kept by default — most have none.
|
||||
|
||||
**Library Recovery**
|
||||
`arrs_failed_stalled_recovery.sh` — detect and recover failed imports and stalled
|
||||
downloads across all arrs. Blocklists the bad release and triggers a re-search — hands-
|
||||
free overnight recovery.
|
||||
|
||||
**Library Enrichment**
|
||||
`lidarr_missing_art.sh` — fetch missing album and artist artwork from fanart.tv and
|
||||
fallback sources. Never overwrites existing files.
|
||||
|
||||
**Upgrade Propagation**
|
||||
`start_webhook_listener.sh` — Node.js HTTP server that receives Sonarr/Radarr/Lidarr
|
||||
OnUpgrade webhooks. Continuous; started at array start. Writes to
|
||||
`/var/log/varaverk/upgrade_webhook.log`.
|
||||
`upgrade_webhook_handler.sh` — triggered by the webhook listener. Pushes the upgraded
|
||||
item folder to every remote node immediately, then triggers an arr library rescan on
|
||||
each remote so the upgraded file is accepted without triggering a redundant quality search.
|
||||
|
||||
**Discovery**
|
||||
`playback_aware_lidarr_discovery.sh` — behavior-driven music discovery. Scores your
|
||||
Emby play history, runs Last.fm getSimilar on top artists, adds the best matches to
|
||||
Lidarr. 0–5 meaningful adds per week.
|
||||
`playback_aware_radarr_discovery.sh` — behavior-driven movie discovery. Scores recently
|
||||
watched movies, runs TMDB recommendations on seeds, adds top candidates to Radarr.
|
||||
`playback_aware_sonarr_discovery.sh` — behavior-driven TV discovery. Scores recently
|
||||
watched series weighted by user diversity, runs TMDB TV recommendations on seeds, adds
|
||||
top shows to Sonarr. Multi-user design: one person binge-watching does not dominate seeds.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ EXECUTION ORDER ━━━
|
||||
|
||||
**Daily via `daily_sync_maintenance.sh` (DAILY_MAINTENANCE_SCRIPTS):**
|
||||
|
||||
```
|
||||
1. lidarr_release_fixer.sh — correct wrong release editions before cleanup sees them
|
||||
2. lidarr_cleanup.sh — orphan removal (music)
|
||||
3. sonarr_cleanup.sh — orphan removal (TV)
|
||||
4. radarr_cleanup.sh — orphan removal (movies)
|
||||
5. lidarr_missing_art.sh — fetch missing artwork (HOST1 only)
|
||||
6. radarr_tmdb_removed.sh — remove TMDb-dropped movies
|
||||
7. sonarr_tvdb_removed.sh — remove TVDB-dropped series
|
||||
```
|
||||
|
||||
Note: `media_shares_permissions.sh` and `media_cleaner.sh` run before these from
|
||||
`Media/` — permissions and junk removal must complete first.
|
||||
|
||||
**Every 30 min + 4hr via orchestrators (CRITICAL/INTERMEDIATE_MAINTENANCE_SCRIPTS):**
|
||||
|
||||
```
|
||||
arrs_failed_stalled_recovery.sh — failed/stalled queue recovery
|
||||
arr_sync.sh — library sync across all nodes
|
||||
```
|
||||
|
||||
**Weekly via `weekly_sync_maintenance.sh` (WEEKLY_MAINTENANCE_SCRIPTS):**
|
||||
|
||||
```
|
||||
playback_aware_lidarr_discovery.sh — score play history → Last.fm similar → Lidarr
|
||||
playback_aware_radarr_discovery.sh — score watch history → TMDB recommendations → Radarr
|
||||
playback_aware_sonarr_discovery.sh — score episode history → TMDB TV → Sonarr
|
||||
```
|
||||
|
||||
**Continuous (started by `array_started.sh`):**
|
||||
|
||||
```
|
||||
start_webhook_listener.sh — Node.js webhook server; dispatches upgrade_webhook_handler.sh
|
||||
```
|
||||
|
||||
**On every arr upgrade (triggered by webhook):**
|
||||
|
||||
```
|
||||
upgrade_webhook_handler.sh — push upgraded folder to all remote nodes + trigger arr rescan
|
||||
```
|
||||
|
||||
**Why release fixer before cleanup:** the fixer corrects Lidarr's release selection so
|
||||
files get imported. If cleanup ran first, a correctable album could accumulate age toward
|
||||
the orphan threshold before the fixer had a chance to fix it.
|
||||
|
||||
**Why arr_sync before rsync:** once arrs agree on what to track, rsync spreads the actual
|
||||
files. An upgrade on one node — new tracked path, old path no longer in API — gets
|
||||
cleaned by arr_cleanup on all nodes after the next sync cycle.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ HOST AWARENESS ━━━
|
||||
|
||||
Scripts run on both servers via `detect_hosts()`, which aliases all `HOST*_` prefixed vars
|
||||
to their unprefixed names at runtime. No manual `HOST1`/`HOST2` comparisons exist in any
|
||||
script.
|
||||
|
||||
`arr_sync.sh` keeps all arr databases in bidirectional union — either server can download
|
||||
to any share. Arr cleanup uses the union model: a file is only an orphan if the arr on
|
||||
this host doesn't have it indexed. Arr scripts check the aliased URL — if empty (arr not
|
||||
configured on this host), they exit cleanly with no action.
|
||||
|
||||
`lidarr_release_fixer.sh` and `lidarr_missing_art.sh` exit cleanly on hosts without
|
||||
Lidarr configured — no HOST1_LIDARR_URL means nothing runs.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SCRIPTS IN THIS FOLDER ━━━
|
||||
|
||||
| Script | Role | When It Runs |
|
||||
|--------|------|--------------|
|
||||
| `arr_sync.sh` | Full-mesh arr library sync — all nodes track the same content | Every 4hr + weekly before rsync |
|
||||
| `lidarr_release_fixer.sh` | Fix wrong MusicBrainz release editions so files get imported | Daily before lidarr_cleanup |
|
||||
| `lidarr_cleanup.sh` | Delete orphaned music files not tracked by Lidarr | Daily |
|
||||
| `sonarr_cleanup.sh` | Delete orphaned TV files not tracked by Sonarr | Daily |
|
||||
| `radarr_cleanup.sh` | Delete orphaned movie files not tracked by Radarr | Daily |
|
||||
| `arrs_failed_stalled_recovery.sh` | Auto-recover failed imports and stalled downloads | Every 30 min / daily |
|
||||
| `lidarr_missing_art.sh` | Fetch missing album and artist artwork | Daily (HOST1 only) |
|
||||
| `radarr_tmdb_removed.sh` | Remove movies dropped from TMDb | Daily |
|
||||
| `sonarr_tvdb_removed.sh` | Remove series dropped from TVDB | Daily |
|
||||
| `start_webhook_listener.sh` | Node.js webhook server — receive arr OnUpgrade and dispatch handler | Continuous |
|
||||
| `upgrade_webhook_handler.sh` | Push upgraded item folder to remote nodes + trigger arr rescan | On each arr upgrade |
|
||||
| `playback_aware_lidarr_discovery.sh` | Behavior-driven music discovery — Emby plays → Last.fm similar → Lidarr | Weekly |
|
||||
| `playback_aware_radarr_discovery.sh` | Behavior-driven movie discovery — Emby watches → TMDB recommendations → Radarr | Weekly |
|
||||
| `playback_aware_sonarr_discovery.sh` | Behavior-driven TV discovery — Emby episodes → TMDB TV recommendations → Sonarr | Weekly |
|
||||
| `arr_download_orphan_cleaner.sh` | Clear orphaned completed downloads out of the SABnzbd Completed folders | Daily |
|
||||
| `sonarr_classification_scan.sh` | Detect series sitting in the wrong root (anime / kids / general); `--move` acts | Daily |
|
||||
| `radarr_classification_scan.sh` | Same for movies, plus junk-metadata detection via `--remove-junk` | Daily |
|
||||
| `lidarr_duplicate_artist_cleanup.sh` | Remove phantom zero-file duplicate artists; flag real ones for review | Daily |
|
||||
| `arr_cache_prefill.sh` | Warm the shared tracked-data cache so consumers never read cold | Array start + every 4hr |
|
||||
| `arr_corruption_scan.sh` | ffprobe every tracked video for corrupt headers; `--remediate` deletes + re-searches | Weekly |
|
||||
| `arr_full_rescan.sh` | Force a real disk↔database reconciliation on all three arrs | Weekly |
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE NEWER LAYERS ━━━
|
||||
|
||||
The original folder was "delete what the arrs no longer track". These were added as distinct
|
||||
failure modes surfaced — each exists because something went wrong that the cleanups could not
|
||||
have caught.
|
||||
|
||||
### 🗑️ Download-Side Orphans — `arr_download_orphan_cleaner.sh`
|
||||
|
||||
Every cleanup script here walks the **library** side. Nothing walked the **download** side —
|
||||
so completed downloads the arrs had stopped tracking accumulated in SABnzbd's Completed
|
||||
folders indefinitely. Discovered as **755 GB of orphaned TV downloads, oldest from 2022**,
|
||||
filling the cache pool to 89%.
|
||||
|
||||
Classifies every entry as TRACKED / RECENT / JUNK / REDUNDANT / IMPORTABLE / UNMATCHED and
|
||||
acts only on the ones it can justify. The queue is a hard gate: if it cannot be read, the arr
|
||||
is skipped entirely, because without it there is no way to tell an active import from an
|
||||
orphan. A run total over `DOWNLOAD_ORPHAN_MAX_DELETE_GB` aborts — an abnormally large delete
|
||||
is the visible symptom of a partial queue fetch.
|
||||
|
||||
### 🎭 Wrong-Root Detection — `sonarr_classification_scan.sh` + `radarr_classification_scan.sh`
|
||||
|
||||
Overseerr lets any user request content into the wrong root folder — kids shows into general
|
||||
TV, anime into Kids_Tv_Shows. These classify every item from metadata alone (genre,
|
||||
certification, network/studio, original language) and report where the computed classification
|
||||
disagrees with the folder the item actually sits in.
|
||||
|
||||
Report-only by default. `--move` acts on forward misplacements and adult-content-in-kids-root
|
||||
leaks. It deliberately does **not** move non-anime content out of the anime root — deliberate
|
||||
style placements (Western animation grouped with anime by choice) are genuine judgment calls.
|
||||
|
||||
Both poll the arr's async move command to completion before verifying, because `moveFiles=true`
|
||||
flips the database instantly while the physical move is still queued behind others.
|
||||
|
||||
### 🎨 Phantom Artists — `lidarr_duplicate_artist_cleanup.sh`
|
||||
|
||||
MusicBrainz duplicates leave two Lidarr entries for one artist, one holding the files and one
|
||||
holding nothing. Removes only the zero-file side, with `deleteFiles=false` so nothing on disk
|
||||
is touched. Pairs where both sides hold files are flagged for review, never auto-resolved.
|
||||
|
||||
Gated on the tracked-count floor shared with `lidarr_cleanup.sh` — during a library-wide desync
|
||||
both sides of a real duplicate can read as zero-file phantoms.
|
||||
|
||||
### 🩺 Corruption + Reconciliation — `arr_corruption_scan.sh` + `arr_full_rescan.sh`
|
||||
|
||||
`arr_corruption_scan.sh` ffprobes tracked video files for corrupt headers. Report-only unless
|
||||
`--remediate`, which deletes the file record and triggers an explicit re-search. Requires
|
||||
repeat detections across separate runs before acting, so a transient probe failure cannot
|
||||
delete a healthy file.
|
||||
|
||||
`arr_full_rescan.sh` forces a genuine disk↔database reconciliation. Organic scans only touch
|
||||
files involved in an import, so an untouched library silently drifts — confirmed when Lidarr
|
||||
reported **~23% of its true track count** for 1,004 of 1,357 artists with no scan running and
|
||||
every file present on disk.
|
||||
|
||||
### ⚡ Cache Warmth — `arr_cache_prefill.sh`
|
||||
|
||||
Populates the shared tracked-data cache at array start and every 4 hours, so consumers never
|
||||
pay a cold fetch. Pure enhancement: nothing depends on it having run, and every consumer still
|
||||
writes through on a cold cache.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ HOW THE SCRIPTS RELATE ━━━
|
||||
|
||||
```
|
||||
Every 4hr / weekly (arr sync before rsync):
|
||||
arr_sync.sh ──────────────── syncs tracked IDs across all nodes
|
||||
│ union model: any node adds → all nodes get it
|
||||
│
|
||||
└── then rsync spreads the actual files to all nodes
|
||||
└── then arr_cleanup removes orphans on all nodes (old paths, removed content)
|
||||
|
||||
Daily maintenance window:
|
||||
[Media/media_shares_permissions.sh + media_cleaner.sh run first — from Media/]
|
||||
│
|
||||
▼
|
||||
lidarr_release_fixer.sh ────── reads MBID tag → switches release in Lidarr → RefreshArtist
|
||||
│ (corrected albums get imported before cleanup scans for orphans)
|
||||
▼
|
||||
lidarr_cleanup.sh ──────────── queries Lidarr API → walks /Music → deletes orphans
|
||||
sonarr_cleanup.sh ──────────── queries Sonarr API → walks /Tv_Shows → deletes orphans
|
||||
radarr_cleanup.sh ──────────── queries Radarr API → walks /Movies → deletes orphans
|
||||
│
|
||||
└── each cleanup → notify_emby_scan() → Emby removes ghost entries
|
||||
|
||||
Daily recovery:
|
||||
arrs_failed_stalled_recovery.sh ── importFailed/stalled → blocklist → re-search
|
||||
|
||||
Weekly discovery (WEEKLY_MAINTENANCE_SCRIPTS):
|
||||
playback_aware_lidarr_discovery.sh ─ Emby plays → Last.fm similar → top candidates → Lidarr
|
||||
playback_aware_radarr_discovery.sh ─ Emby watches → TMDB recommendations → top candidates → Radarr
|
||||
playback_aware_sonarr_discovery.sh ─ Emby episodes → TMDB TV recommendations → top candidates → Sonarr
|
||||
│
|
||||
└── each discovery script fires arr search immediately after successful add
|
||||
|
||||
Continuous (started by array_started.sh):
|
||||
start_webhook_listener.sh ── Node.js HTTP server listens on WEBHOOK_PORT
|
||||
│ arr OnUpgrade fires webhook → POST to http://HOST_LAN_IP:WEBHOOK_PORT/webhook?key=SECRET
|
||||
└── upgrade_webhook_handler.sh
|
||||
├── rsync upgraded folder → all remote nodes immediately
|
||||
└── trigger arr library rescan on each remote (accept new file, no quality search)
|
||||
|
||||
Ad-hoc enrichment:
|
||||
lidarr_missing_art.sh ─────── discovers missing artwork → fetches from fanart.tv
|
||||
radarr_tmdb_removed.sh ────── status="deleted" → remove from Radarr + add exclusion
|
||||
sonarr_tvdb_removed.sh ────── status="deleted" → remove from Sonarr + add exclusion
|
||||
```
|
||||
Executable
+182
@@ -0,0 +1,182 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Arr Cache Prefill ===========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Populates the shared tracked-data cache (see arr_get_tracked_data() in common.sh) for
|
||||
# Lidarr, Sonarr, and Radarr. Without this, each arr's cache stays cold until whichever
|
||||
# script happens to touch that arr first writes through — which could be hours, depending
|
||||
# on the daily schedule. Originally Lidarr-only (lidarr_cache_prefill.sh, 2026-07-16),
|
||||
# generalized the same day to cover all three arrs once the cache mechanism itself was
|
||||
# generalized.
|
||||
#
|
||||
# Runs on two schedules (2026-07-17): once at array start (closes the cold-boot gap, 10min
|
||||
# wait ceiling for slow-starting containers) and again every 30min via
|
||||
# CRITICAL_MAINTENANCE_SCRIPTS with a 1min wait ceiling (ARR_PREFILL_WAIT_MINUTES=1 override)
|
||||
# — a live fetch+write takes seconds, so there's no reason to tolerate the boot-time wait on
|
||||
# a recurring job. This is what keeps the cache-first consumers' data reliably under 30min
|
||||
# old instead of only refreshing whenever some other script happens to write through.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Each arr's container may still be starting when this fires (array just started) — retries
|
||||
# reaching that arr's API for up to ARR_PREFILL_WAIT_MINUTES before giving up on it and moving
|
||||
# to the next. Not fatal if one never comes up in time; that arr's cache just stays cold until
|
||||
# the next script writes through naturally, exactly as it would without this script existing.
|
||||
# An arr not configured on this host (e.g. Lidarr is HOST1-only) is skipped cleanly.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Warm the Cache, Never Own It
|
||||
# This script only pre-populates what arr_get_tracked_data() would fetch on demand
|
||||
# anyway. Nothing depends on it having run — every consumer still writes through on a
|
||||
# cold cache. It removes latency and staleness, it is not a dependency.
|
||||
#
|
||||
# Failure Is a No-Op, Not an Error
|
||||
# An arr that never comes up in time simply leaves its cache cold, exactly as if this
|
||||
# script did not exist. That is why a missed prefill is logged rather than notified —
|
||||
# the fallback path is the normal path.
|
||||
#
|
||||
# Wait Ceiling Matched to the Trigger
|
||||
# The array-start run tolerates a long wait because containers are genuinely still
|
||||
# starting. The 30-minute recurring run does not, because a live fetch takes seconds
|
||||
# and a long wait there would only serve to overlap the next tick.
|
||||
#
|
||||
# Per-Arr Independence
|
||||
# Each arr is prefilled on its own. One unconfigured or slow-starting arr never
|
||||
# prevents the other two from being warmed.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root required — chown-free here, but matches convention across Arrs_Stack/
|
||||
# jq required — skips the whole run cleanly (not fatal) if jq is missing
|
||||
# acquire_lock — prevents two invocations of this script overlapping (default strict
|
||||
# mode) — matters if the boot-time run is still waiting on a slow
|
||||
# container when the first 30min critical-tier tick fires
|
||||
# Reachability retry — tolerates a slow-starting container up to ARR_PREFILL_WAIT_MINUTES;
|
||||
# never fatal if one never comes up, that arr's cache just stays cold
|
||||
# Active-rescan check — skips an arr this cycle if a rescan-type command is running, instead
|
||||
# of a live fetch arr_cache_write() would refuse to persist anyway
|
||||
# (2026-07-17) — avoids wasted API calls during a long rescan
|
||||
# Per-arr isolation — one arr failing or timing out never blocks or fails the others
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY
|
||||
# HOST*_SONARR_URL / HOST*_SONARR_API_KEY
|
||||
# HOST*_RADARR_URL / HOST*_RADARR_API_KEY
|
||||
# All aliased by detect_hosts()
|
||||
#
|
||||
# master.conf
|
||||
# ARR_PREFILL_WAIT_MINUTES — how long to retry reaching each arr before giving up on it
|
||||
# (default 10). The CRITICAL_MAINTENANCE_SCRIPTS entry overrides this to 1 via parse_args'
|
||||
# VAR=VAL mechanism for the 30min recurring run — the 10min default is sized for cold boot,
|
||||
# not a job that fires every half hour.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arr_cache_prefill.sh
|
||||
# Normal run — populates all three arr caches, or skips cleanly per-arr as described
|
||||
# under OPERATIONAL SAFEGUARDS above. No --dry-run/--status mode: this script only ever
|
||||
# reads and writes cache, there's no destructive action to preview and no separate state
|
||||
# worth inspecting beyond the cache files themselves (see Tools/arr_rescan_monitor.sh
|
||||
# --status for cache age/active-rescan inspection).
|
||||
#
|
||||
# arr_cache_prefill.sh --log
|
||||
# Verbose — per-arr detail as each one is checked/fetched/skipped.
|
||||
#
|
||||
# arr_cache_prefill.sh ARR_PREFILL_WAIT_MINUTES=1
|
||||
# Override the reachability-retry ceiling for this run only (parse_args VAR=VAL
|
||||
# mechanism) — this is how CRITICAL_MAINTENANCE_SCRIPTS invokes it every 30min.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
warn "jq not found — skipping arr cache prefill"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
WAIT_MINUTES="${ARR_PREFILL_WAIT_MINUTES:-10}"
|
||||
|
||||
# Args: arr_type, url, api_key, api_version
|
||||
_prefill_one() {
|
||||
local arr_type="$1" url="$2" api_key="$3" api_version="$4"
|
||||
|
||||
if [[ -z "$url" ]] || [[ -z "$api_key" ]]; then
|
||||
info "${arr_type^} not configured on $MY_ID ($LOCAL_SERVER_NAME) — nothing to prefill"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local waited=0
|
||||
until check_api "$url" "${arr_type^}" 5 >/dev/null 2>&1; do
|
||||
if [[ "$waited" -ge $(( WAIT_MINUTES * 60 )) ]]; then
|
||||
warn "${arr_type^} not reachable after ${WAIT_MINUTES}m — leaving cache cold, next script will write through"
|
||||
return 0
|
||||
fi
|
||||
sleep 15
|
||||
(( waited += 15 ))
|
||||
done
|
||||
|
||||
# Skip cleanly if a rescan-type command is already active — arr_cache_write() below
|
||||
# would refuse to persist a mid-rescan snapshot anyway (2026-07-17 guard), so fetching
|
||||
# it live first would just be a wasted API call every time this fires during a long
|
||||
# rescan. Matches the same check arr_full_rescan.sh and the cleanup scripts already do.
|
||||
local active_cmd
|
||||
active_cmd=$(arr_active_rescan_command "$arr_type" "$url" "$api_key" "$api_version")
|
||||
if [[ -n "$active_cmd" ]]; then
|
||||
info "${arr_type^} mid-rescan ($active_cmd) — skipping this cycle, cache stays as-is"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local endpoint="${ARR_LIBRARY_ENDPOINT[$arr_type]:-}"
|
||||
if [[ -z "$endpoint" ]]; then
|
||||
warn "No library endpoint known for ${arr_type} — skipping"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local items
|
||||
items=$(arr_api "$url" "$api_key" "$api_version" "$endpoint" "${arr_type^}") || {
|
||||
warn "Could not fetch ${arr_type} library for cache prefill — leaving cache cold"
|
||||
return 0
|
||||
}
|
||||
|
||||
if arr_cache_write "$arr_type" "$items"; then
|
||||
log "$ICON_DONE ${arr_type^} cache prefilled ($(echo "$items" | jq 'length') items)"
|
||||
else
|
||||
warn "Failed to write ${arr_type} cache prefill"
|
||||
fi
|
||||
}
|
||||
|
||||
_prefill_one "lidarr" "${LIDARR_URL:-}" "${LIDARR_API_KEY:-}" "v1"
|
||||
_prefill_one "sonarr" "${SONARR_URL:-}" "${SONARR_API_KEY:-}" "v3"
|
||||
_prefill_one "radarr" "${RADARR_URL:-}" "${RADARR_API_KEY:-}" "v3"
|
||||
|
||||
exit 0
|
||||
@@ -0,0 +1,789 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================ Arr Corruption Scan ==============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Scans Sonarr's and Radarr's tracked video files for corrupt headers (ffprobe-based, same
|
||||
# detection method as the third-party Healarr tool) and, in --remediate mode, deletes the bad
|
||||
# file from the owning arr and explicitly triggers a search to replace it.
|
||||
#
|
||||
# Built after Healarr crashed mid-scan on a genuine Go concurrency bug (unsynchronized
|
||||
# map access when multiple corruption events land at once — confirmed via its own crash
|
||||
# log, not fixable from our side). The core idea (scan → delete → re-search) isn't hard to
|
||||
# replicate; the fix here is architectural: this script processes one file at a time,
|
||||
# strictly sequential, so the race condition that killed Healarr can't happen — there's
|
||||
# nothing running concurrently to race.
|
||||
#
|
||||
# Sonarr-only originally (2026-07-18/19); Radarr/Movies coverage added 2026-07-21 as a second
|
||||
# arr in the same per-file scan/strike/remediate loop, not a separate script — the detection,
|
||||
# strike, and state-file logic is identical, only the API shape (episodefile vs moviefile,
|
||||
# EpisodeSearch vs MoviesSearch) differs. Radarr's moviefile list is fetched batched
|
||||
# (movieId=... query params, BATCH_SIZE at a time) rather than off the movie list's embedded
|
||||
# .movieFile alone — Radarr supports a second tracked file per movie (alternate editions/
|
||||
# extras) that never shows up there, same gap radarr_cleanup.sh hit and fixed 2026-07-19;
|
||||
# reusing that batched-fetch shape here instead of the simpler single-file read so a
|
||||
# corruption scan doesn't silently skip every alternate edition in the library.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# WHY A SEPARATE CONTAINER FOR FFPROBE
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Neither Sonarr nor Radarr bundle ffprobe. ffprobe runs via `docker exec` into a
|
||||
# different container that does — confirmed live 2026-07-18:
|
||||
# Jellyfin — working ffprobe, mounts every share Emby does (Tv_Shows, Movies, kids/
|
||||
# anime shares, standup) as of 2026-07-18
|
||||
# Emby — mounts everything too, but its bundled ffprobe binary is broken
|
||||
# (2017-dated, fails to exec — likely a missing dynamic linker
|
||||
# dependency, not something to fix here)
|
||||
# Jellyfin is what's configured (HOST*_FFPROBE_CONTAINER).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Runs Sonarr then Radarr, sequentially, never parallel — one arr failing/unconfigured never
|
||||
# blocks the other. Within each arr, one file at a time, in this order per file:
|
||||
# 1. Skip if unchanged (mtime+size) since the last time it verified clean — state file
|
||||
# avoids re-probing the entire library every run, which would take far too long at
|
||||
# this library size (90k+ tracked files).
|
||||
# 2. ffprobe via `docker exec` into FFPROBE_CONTAINER. Empty stderr + exit 0 = clean.
|
||||
# Anything else = corrupt (same signature as Healarr: "Invalid data found when
|
||||
# processing input", EBML header errors, etc.)
|
||||
# 3. Report-only by default. --remediate additionally:
|
||||
# a. DELETE the specific episodefile/moviefile record via the arr's API
|
||||
# b. Verify hasFile flipped false (never trust the DELETE response alone)
|
||||
# c. Explicitly trigger EpisodeSearch/MoviesSearch for that episode/movie — this is
|
||||
# deliberate, not left to the arr's own background missing-search cycle, because
|
||||
# that cycle skips unmonitored items entirely. An explicit search call does not
|
||||
# have that restriction (confirmed live: two unmonitored episodes Healarr healed
|
||||
# both still got successfully re-grabbed via this exact same kind of search call,
|
||||
# logged in Sonarr's history as "UserInvokedSearch").
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Detect Always, Act Only on Request
|
||||
# A bare run probes and reports. Deleting a file the arr believes it has is a
|
||||
# destructive act, so it requires --remediate explicitly. The scan can be scheduled
|
||||
# weekly and read without any risk of it removing media on its own.
|
||||
#
|
||||
# One Probe Result Is Not Evidence
|
||||
# ffprobe can fail for reasons that have nothing to do with the file — a mid-write
|
||||
# import, an NFS blip, a container restart. Corruption must be observed
|
||||
# CORRUPTION_SCAN_STRIKE_LIMIT times consecutively before remediation acts, and a
|
||||
# single clean re-probe resets the counter.
|
||||
#
|
||||
# Skip-Cache Over Re-Probing
|
||||
# At 90k+ tracked files a full re-probe every run is not viable. Files unchanged by
|
||||
# mtime and size since they last verified clean are skipped, so each run spends its
|
||||
# time on what actually changed rather than re-proving the library from scratch.
|
||||
#
|
||||
# Delete the Record, Let the Arr Re-Acquire
|
||||
# Remediation removes the file record and explicitly triggers a search. The arr is
|
||||
# left to obtain a good copy through its normal path — this script never tries to
|
||||
# repair a file in place.
|
||||
#
|
||||
# Explicit Search, Not the Background Cycle
|
||||
# The re-search is triggered directly rather than left to the arr's own missing-search
|
||||
# cycle, because that cycle skips unmonitored items entirely and would silently leave
|
||||
# an unmonitored corrupt file deleted and never replaced.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# docker exec into the ffprobe container requires root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents overlapping runs. Two instances would both probe and could
|
||||
# both count a strike against the same file, reaching the limit in half the intended
|
||||
# number of observations.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases the arr URLs, API keys and HOST*_FFPROBE_CONTAINER.
|
||||
#
|
||||
# jq Dependency Check
|
||||
# Fails fast if jq is missing — the tracked-file lists and every hasFile verification
|
||||
# are parsed with it.
|
||||
#
|
||||
# Report-Only Default
|
||||
# Nothing is deleted without --remediate.
|
||||
#
|
||||
# ffprobe Configuration Check
|
||||
# Exits cleanly if FFPROBE_CONTAINER / FFPROBE_BIN are unconfigured for this host.
|
||||
#
|
||||
# ffprobe Container Health Check
|
||||
# check_container_health() verifies the container is running and healthy before any
|
||||
# probing. Every probe is a docker exec into it — if it is stopped or unhealthy every
|
||||
# exec fails, every file reads as corrupt, and two consecutive runs would clear the
|
||||
# strike limit and hand --remediate the whole library to delete.
|
||||
#
|
||||
# API Reachability + Version Gate
|
||||
# check_api then check_arr_version per arr. A version mismatch skips that arr rather
|
||||
# than issuing deletes against an API whose file-record endpoints may have moved.
|
||||
#
|
||||
# Per-Arr Isolation
|
||||
# Sonarr and Radarr run sequentially, and one failing, unconfigured or version-
|
||||
# mismatched arr never blocks the other.
|
||||
#
|
||||
# Unmapped Path Skip
|
||||
# Files whose arr-side path cannot be mapped into the ffprobe container's mount
|
||||
# namespace are skipped and counted, never probed through a wrong path and never
|
||||
# treated as corrupt because the probe could not see them.
|
||||
#
|
||||
# Strike Threshold
|
||||
# CORRUPTION_SCAN_STRIKE_LIMIT consecutive corrupt detections are required before
|
||||
# --remediate deletes anything. A transient ffprobe failure cannot trigger a delete,
|
||||
# and a clean re-probe clears the counter.
|
||||
#
|
||||
# Post-Delete Verification
|
||||
# The DELETE response is never trusted. hasFile is re-checked and must have flipped
|
||||
# false before the re-search is issued, so a failed delete never leaves the arr
|
||||
# searching for something it still believes it has.
|
||||
#
|
||||
# Targeted Deletion
|
||||
# Only the specific episodefile/moviefile record for the corrupt file is removed —
|
||||
# never the series, movie, or any sibling file.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
# SONARR_URL / SONARR_API_KEY, RADARR_URL / RADARR_API_KEY — existing, aliased by
|
||||
# detect_hosts(). Either arr missing its URL/key is skipped, not fatal.
|
||||
# HOST*_FFPROBE_CONTAINER — container name with a working ffprobe binary
|
||||
# HOST*_FFPROBE_BIN — full path to that binary inside the container
|
||||
# HOST*_FFPROBE_PATH_MAP — host path prefix → that container's internal path prefix
|
||||
# (separate from the arrs' own path maps — the ffprobe
|
||||
# container almost certainly mounts shares differently)
|
||||
#
|
||||
# master.conf
|
||||
# CORRUPTION_SCAN_STATE_FILE — path to the clean-file skip-cache (default in DATA_DIR),
|
||||
# shared across both arrs — keyed by host path, which never
|
||||
# collides between a Sonarr and a Radarr share
|
||||
# CORRUPTION_SCAN_STRIKES_FILE — path to the consecutive-corrupt-detection counter (default
|
||||
# in DATA_DIR), keyed by host path
|
||||
# CORRUPTION_SCAN_STRIKE_LIMIT — consecutive corrupt detections required before --remediate
|
||||
# acts on a file (default 2) — guards against a one-off
|
||||
# ffprobe hiccup (mid-write file, NFS blip) triggering an
|
||||
# unnecessary delete+re-search. Resets on a clean re-probe.
|
||||
# SONARR_VERSION_MAJOR / RADARR_VERSION_MAJOR — reused from sonarr_cleanup.sh/
|
||||
# radarr_cleanup.sh for the API version check
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arr_corruption_scan.sh — report-only, scans everything not yet
|
||||
# verified clean (Sonarr then Radarr)
|
||||
# arr_corruption_scan.sh --remediate — delete + re-search on every corrupt file found
|
||||
# arr_corruption_scan.sh --limit=50 — cap EACH arr to 50 newly-probed files this run
|
||||
# (state file makes repeat runs cheap regardless,
|
||||
# but useful for a bounded first test)
|
||||
# arr_corruption_scan.sh --log — verbose (prints every clean file too)
|
||||
# arr_corruption_scan.sh --status — show config and exit
|
||||
# arr_corruption_scan.sh --filter=Becker — only consider paths containing this substring
|
||||
# (testing/targeting a specific show/movie; state
|
||||
# file and everything else behaves normally)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# --remediate / --limit are script-local, not recognized by parse_args — check raw args
|
||||
# before they get filtered.
|
||||
REMEDIATE=false
|
||||
SCAN_LIMIT=0
|
||||
PATH_FILTER=""
|
||||
for _arg in "$@"; do
|
||||
case "$_arg" in
|
||||
--remediate) REMEDIATE=true ;;
|
||||
--limit=*) SCAN_LIMIT="${_arg#*=}" ;;
|
||||
--filter=*) PATH_FILTER="${_arg#*=}" ;;
|
||||
esac
|
||||
done
|
||||
unset _arg
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for arr API calls"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock "wait"
|
||||
TMP_DIR="/tmp/arr_corruption_scan_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "_release_all_locks; rm -rf $TMP_DIR" EXIT
|
||||
|
||||
detect_hosts
|
||||
|
||||
# FFPROBE_* aren't part of the shared arr alias set in detect_hosts() — resolve them here,
|
||||
# same eval-based pattern build_arr_path_map() uses for the associative array.
|
||||
FFPROBE_CONTAINER_VAR="${MY_ID}_FFPROBE_CONTAINER"
|
||||
FFPROBE_CONTAINER="${!FFPROBE_CONTAINER_VAR:-}"
|
||||
FFPROBE_BIN_VAR="${MY_ID}_FFPROBE_BIN"
|
||||
FFPROBE_BIN="${!FFPROBE_BIN_VAR:-}"
|
||||
|
||||
declare -A FFPROBE_PATH_MAP=()
|
||||
_fp_map_var="${MY_ID}_FFPROBE_PATH_MAP"
|
||||
eval "for key in \"\${!${_fp_map_var}[@]}\"; do
|
||||
FFPROBE_PATH_MAP[\"\$key\"]=\"\${${_fp_map_var}[\$key]}\"
|
||||
done"
|
||||
unset _fp_map_var
|
||||
|
||||
if [[ -z "$FFPROBE_CONTAINER" || -z "$FFPROBE_BIN" ]]; then
|
||||
error "FFPROBE_CONTAINER/FFPROBE_BIN not configured on $MY_ID — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Every probe is a docker exec into this container. If it is stopped or unhealthy, every
|
||||
# exec fails, every file reads as corrupt, and two such runs would clear the strike limit
|
||||
# and hand --remediate an entire library to delete. Abort before probing anything.
|
||||
check_container_health "$FFPROBE_CONTAINER" "${DOCKER_TIMEOUT:-30}" "Arr Corruption Scan"
|
||||
|
||||
CORRUPTION_SCAN_STATE_FILE="${CORRUPTION_SCAN_STATE_FILE:-$DATA_DIR/corruption_scan_state.tsv}"
|
||||
mkdir -p "$(dirname "$CORRUPTION_SCAN_STATE_FILE")"
|
||||
touch "$CORRUPTION_SCAN_STATE_FILE"
|
||||
|
||||
CORRUPTION_SCAN_STRIKES_FILE="${CORRUPTION_SCAN_STRIKES_FILE:-$DATA_DIR/corruption_scan_strikes.tsv}"
|
||||
CORRUPTION_SCAN_STRIKE_LIMIT="${CORRUPTION_SCAN_STRIKE_LIMIT:-2}"
|
||||
CORRUPTION_SCAN_MAX_CORRUPT_PCT="${CORRUPTION_SCAN_MAX_CORRUPT_PCT:-10}"
|
||||
CORRUPTION_SCAN_MAX_CONSECUTIVE="${CORRUPTION_SCAN_MAX_CONSECUTIVE:-15}"
|
||||
CORRUPTION_SCAN_GUARD_MIN_SCANNED="${CORRUPTION_SCAN_GUARD_MIN_SCANNED:-20}"
|
||||
mkdir -p "$(dirname "$CORRUPTION_SCAN_STRIKES_FILE")"
|
||||
touch "$CORRUPTION_SCAN_STRIKES_FILE"
|
||||
|
||||
# Thin wrappers around common.sh's wd_state_get/wd_state_set — same shape as
|
||||
# stability_watchdog.sh's get_strikes/set_strikes/increment_strikes/reset_strikes, keyed here
|
||||
# by host path instead of a watchdog check name. Requires repeat corrupt detections across
|
||||
# separate scan runs before --remediate acts, so a one-off ffprobe hiccup (mid-write file,
|
||||
# NFS blip) can't trigger an unnecessary delete+re-search on its own.
|
||||
get_scan_strikes() {
|
||||
wd_state_get "$1" "$CORRUPTION_SCAN_STRIKES_FILE"
|
||||
}
|
||||
set_scan_strikes() {
|
||||
wd_state_set "$1" "$2" "$CORRUPTION_SCAN_STRIKES_FILE"
|
||||
}
|
||||
increment_scan_strikes() {
|
||||
local current
|
||||
current=$(get_scan_strikes "$1")
|
||||
[[ -z "$current" ]] && current=0
|
||||
(( current++ ))
|
||||
set_scan_strikes "$1" "$current"
|
||||
echo "$current"
|
||||
}
|
||||
reset_scan_strikes() {
|
||||
local current
|
||||
current=$(get_scan_strikes "$1")
|
||||
[[ -n "$current" && "$current" != "0" ]] && set_scan_strikes "$1" 0
|
||||
}
|
||||
|
||||
# Bails out of the whole run without committing anything. Safe to call at any point before
|
||||
# the commit phase: strikes are queued in memory until then, so an abort leaves the strike
|
||||
# file exactly as the previous run left it and deletes nothing.
|
||||
abort_scan() {
|
||||
local why="$1"
|
||||
error "Corruption scan ABORTED — $why"
|
||||
error "No strikes recorded and nothing remediated this run — the library was not trusted."
|
||||
[[ -n "${FRESH_CLEAN_TMP:-}" ]] && rm -f "$FRESH_CLEAN_TMP"
|
||||
notify "Corruption scan aborted on $(hostname) ($MY_ID) — $why. Nothing deleted." \
|
||||
"Arr Corruption Scan" "warning"
|
||||
exit 1
|
||||
}
|
||||
|
||||
# Per-arr API shape differences — everything else in the scan/strike/remediate loop below is
|
||||
# identical between Sonarr and Radarr.
|
||||
declare -A ARR_FILE_ENDPOINT=( [sonarr]="episodefile" [radarr]="moviefile" )
|
||||
declare -A ARR_PARENT_ENDPOINT=( [sonarr]="episode" [radarr]="movie" )
|
||||
declare -A ARR_SEARCH_COMMAND=( [sonarr]="EpisodeSearch" [radarr]="MoviesSearch" )
|
||||
declare -A ARR_SEARCH_ID_FIELD=( [sonarr]="episodeIds" [radarr]="movieIds" )
|
||||
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
for arr in sonarr radarr; do
|
||||
url_var="${arr^^}_URL"
|
||||
echo "$ICON_GEAR ${arr^} URL: ${!url_var:-not configured}"
|
||||
done
|
||||
echo "$ICON_GEAR FFprobe container: $FFPROBE_CONTAINER"
|
||||
echo "$ICON_GEAR FFprobe binary: $FFPROBE_BIN"
|
||||
echo "$ICON_GEAR FFprobe path map: ${#FFPROBE_PATH_MAP[@]} entries"
|
||||
echo "$ICON_GEAR State file: $CORRUPTION_SCAN_STATE_FILE"
|
||||
echo "$ICON_GEAR Strike limit: $CORRUPTION_SCAN_STRIKE_LIMIT"
|
||||
echo "$ICON_GEAR Remediate: $REMEDIATE"
|
||||
echo "$ICON_GEAR Corrupt ceiling: ${CORRUPTION_SCAN_MAX_CORRUPT_PCT}% of scanned (min ${CORRUPTION_SCAN_GUARD_MIN_SCANNED} scanned)"
|
||||
echo "$ICON_GEAR Consecutive trip: $CORRUPTION_SCAN_MAX_CONSECUTIVE"
|
||||
echo "$ICON_GEAR Scan limit: ${SCAN_LIMIT:-unlimited} (per arr)"
|
||||
echo "$ICON_GEAR Path filter: ${PATH_FILTER:-none}"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo ""
|
||||
[[ "$REMEDIATE" == true ]] && warn "REMEDIATE MODE — corrupt files will be deleted and re-searched" \
|
||||
|| info "Report-only — pass --remediate to act"
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
check_container_health "$FFPROBE_CONTAINER" 15 "Corruption Scan"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Translates a host filesystem path to FFPROBE_CONTAINER's internal path via prefix match
|
||||
# against FFPROBE_PATH_MAP. Empty output (return 1) means this file's share isn't covered
|
||||
# by the ffprobe container yet — caller must skip, not guess.
|
||||
ffprobe_translate_path() {
|
||||
local host_path="$1" prefix
|
||||
for prefix in "${!FFPROBE_PATH_MAP[@]}"; do
|
||||
if [[ "$host_path" == "$prefix"/* ]]; then
|
||||
echo "${FFPROBE_PATH_MAP[$prefix]}${host_path#$prefix}"
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
# Probes one file. Echoes "clean" or "corrupt:<reason>". Never trusts a truncated/garbled
|
||||
# stderr as automatically corrupt — only a real non-empty ffprobe stderr counts.
|
||||
probe_file() {
|
||||
local host_path="$1" container_path output rc
|
||||
container_path=$(ffprobe_translate_path "$host_path") || { echo "unmapped"; return; }
|
||||
output=$(docker exec "$FFPROBE_CONTAINER" "$FFPROBE_BIN" -v error "$container_path" 2>&1)
|
||||
rc=$?
|
||||
|
||||
# docker exec writes its own failures to the same stream ffprobe uses, so a stopped
|
||||
# container or an unreachable daemon is otherwise indistinguishable from a corrupt
|
||||
# header. A stopped container exits 1 with a daemon message; a missing binary exits
|
||||
# 127 — neither is evidence about the file, so both must be caught.
|
||||
if (( rc >= 125 )) \
|
||||
|| [[ "$output" == "Error response from daemon:"* \
|
||||
|| "$output" == "Cannot connect to the Docker daemon"* \
|
||||
|| "$output" == "error during connect:"* ]]; then
|
||||
echo "probe_error:${output//$'\n'/ }"
|
||||
return
|
||||
fi
|
||||
|
||||
if [[ -z "$output" ]]; then
|
||||
echo "clean"
|
||||
elif (( rc != 0 )); then
|
||||
# ffprobe could not parse the file — EBML header parsing failed, moov atom not found,
|
||||
# contradictionary STSC and STCO. This is the only class that may be remediated.
|
||||
echo "corrupt:${output//$'\n'/ }"
|
||||
else
|
||||
# Exit 0 with stderr output: a recoverable muxing complaint, most commonly
|
||||
# "Referenced QT chapter track not found", which many recent .mp4 releases emit and
|
||||
# which says nothing about playability. Equating any stderr with corruption is what
|
||||
# produced 103 "corrupt" files on 2026-08-23 — 28 of 43 newly scanned Radarr items.
|
||||
# Reported for visibility, never strike-tracked, never remediated.
|
||||
echo "suspect:${output//$'\n'/ }"
|
||||
fi
|
||||
}
|
||||
|
||||
# Appends one arr_api() call's output to a batch file, but ONLY on success. arr_api() prints
|
||||
# its own error message via error() (a plain `echo`, i.e. stdout, not stderr) on any non-200
|
||||
# response — appending its raw output unconditionally means a single failed batch call (one
|
||||
# bad seriesId/movieId batch out of hundreds) mixes a plain-text error line into what's
|
||||
# otherwise a stream of valid JSON arrays, and `jq -s` then fails to parse the WHOLE file,
|
||||
# turning one bad batch into zero usable files for the entire arr. Confirmed live 2026-07-21:
|
||||
# Sonarr seriesId=650 returned HTTP 404 (stale/deleted series reference) mid-walk, and that
|
||||
# single 404's error text corrupted the full 138MB/1177-series concatenated batch, silently
|
||||
# zeroing out the whole Sonarr scan for that run. Capturing output first and gating the
|
||||
# append on the actual exit code isolates one bad call to just that call.
|
||||
_arr_api_append_on_success() {
|
||||
local out
|
||||
out=$(arr_api "$1" "$2" "$3" "$4" "$5" 2>/dev/null)
|
||||
[[ $? -eq 0 ]] && echo "$out" >> "$6"
|
||||
}
|
||||
|
||||
# Fetches every movie's file(s) via Radarr's moviefile endpoint, batched (movieId=X repeated
|
||||
# query param, BATCH_SIZE at a time — a single whole-library request 414s, confirmed live by
|
||||
# radarr_cleanup.sh 2026-07-19). Deliberately not just the movie list's embedded .movieFile —
|
||||
# that only ever has the primary file, missing Radarr's second-tracked-file-per-movie feature
|
||||
# (alternate editions/extras). Echoes the raw moviefile JSON array (id/movieId/path per item).
|
||||
fetch_radarr_items() {
|
||||
local movies_now
|
||||
movies_now=$(arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "movie" "Radarr" 2>/dev/null)
|
||||
|
||||
local _ids=() _id _qs="" _batch_count=0
|
||||
local BATCH_SIZE=200 # 250 confirmed working live 2026-07-19 by radarr_cleanup.sh, margin kept
|
||||
mapfile -t _ids < <(echo "$movies_now" | jq -r '.[] | select(.hasFile==true) | .id' 2>/dev/null)
|
||||
|
||||
local all_tmp; all_tmp=$(mktemp)
|
||||
for _id in "${_ids[@]}"; do
|
||||
_qs+="movieId=${_id}&"
|
||||
(( _batch_count++ ))
|
||||
if [[ "$_batch_count" -ge "$BATCH_SIZE" ]]; then
|
||||
_arr_api_append_on_success "$RADARR_URL" "$RADARR_API_KEY" "v3" "moviefile?${_qs%&}" "Radarr" "$all_tmp"
|
||||
_qs=""
|
||||
_batch_count=0
|
||||
fi
|
||||
done
|
||||
if [[ -n "$_qs" ]]; then
|
||||
_arr_api_append_on_success "$RADARR_URL" "$RADARR_API_KEY" "v3" "moviefile?${_qs%&}" "Radarr" "$all_tmp"
|
||||
fi
|
||||
|
||||
jq -s 'add // []' "$all_tmp" 2>/dev/null
|
||||
rm -f "$all_tmp"
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Load clean-file state (skip cache) ━━━
|
||||
# ==============================================================================================
|
||||
declare -A CLEAN_STATE
|
||||
while IFS=$'\t' read -r _s_path _s_stamp; do
|
||||
[[ -n "$_s_path" ]] && CLEAN_STATE["$_s_path"]="$_s_stamp"
|
||||
done < "$CORRUPTION_SCAN_STATE_FILE"
|
||||
unset _s_path _s_stamp
|
||||
info "Loaded ${#CLEAN_STATE[@]} previously-verified-clean entries"
|
||||
|
||||
# Merges a fresh-clean-stamps temp file into the persistent state file, newest wins per path —
|
||||
# reading fresh entries first (before the old base file) means the first occurrence tac/awk
|
||||
# keeps is always the newest one for any path re-verified this run. Called once per arr
|
||||
# (immediately after that arr's scan, not batched to the very end of the whole script) so a
|
||||
# hard-exit partway through the NEXT arr — check_container_health()/check_arr_version() both
|
||||
# exit 1 directly on a real failure, not just return — can never wipe out the previous arr's
|
||||
# already-computed clean state for this run.
|
||||
merge_clean_state() {
|
||||
local fresh_tmp="$1" state_tmp
|
||||
state_tmp=$(mktemp)
|
||||
cat "$fresh_tmp" "$CORRUPTION_SCAN_STATE_FILE" | awk -F'\t' '!seen[$1]++' | sort > "$state_tmp"
|
||||
mv "$state_tmp" "$CORRUPTION_SCAN_STATE_FILE"
|
||||
}
|
||||
|
||||
TOTAL_SCANNED=0
|
||||
TOTAL_CORRUPT=0
|
||||
TOTAL_REMEDIATED=0
|
||||
TOTAL_REMEDIATE_FAILED=0
|
||||
declare -A ARR_SCANNED ARR_SKIPPED_CACHED ARR_SKIPPED_UNMAPPED ARR_CORRUPT ARR_SUSPECT ARR_PROBE_ERRORS ARR_STRIKE_HELD ARR_REMEDIATED ARR_REMEDIATE_FAILED
|
||||
|
||||
for arr in sonarr radarr; do
|
||||
url_var="${arr^^}_URL"; key_var="${arr^^}_API_KEY"
|
||||
# Named arr_url/arr_key, not url/key — build_arr_path_map() below uses a non-local
|
||||
# `for key in ...` loop internally (iterating FFPROBE/path-map prefixes) and would
|
||||
# silently clobber a plain $key with its last loop value otherwise. Confirmed live
|
||||
# 2026-07-21: this exact collision fed a path-map prefix ("/ext-anime-shows") to Sonarr's
|
||||
# API calls as the X-Api-Key header instead of the real key, making every Sonarr call
|
||||
# this loop made fail with 401 while looking like a connectivity problem.
|
||||
arr_url="${!url_var:-}"; arr_key="${!key_var:-}"
|
||||
|
||||
if [[ -z "$arr_url" || -z "$arr_key" ]]; then
|
||||
info "${arr^} not configured on $MY_ID ($LOCAL_SERVER_NAME) — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo " ${arr^} — $arr_url"
|
||||
|
||||
build_arr_path_map "${arr^^}"
|
||||
|
||||
check_container_health "${arr^}" 15 "Corruption Scan"
|
||||
ver_var="${arr^^}_VERSION_MAJOR"
|
||||
check_arr_version "$arr_url" "$arr_key" "v3" "${!ver_var}" "${arr^}" || {
|
||||
warn "${arr^} version check failed — skipping this arr"
|
||||
continue
|
||||
}
|
||||
|
||||
# ──────────────────────────────────────────────────────────────────────────────────────
|
||||
# Fetch tracked files, normalized to {path, file_id, parent_id, title} regardless of arr —
|
||||
# everything past this point is arr-agnostic.
|
||||
# ──────────────────────────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching ${arr^} Tracked Files ━━━"
|
||||
|
||||
if [[ "$arr" == "sonarr" ]]; then
|
||||
# Prefer the shared per-episode-file cache written by sonarr_cleanup.sh (has
|
||||
# id/episodeId/seriesId/path already) — falls back to a live per-series walk only on
|
||||
# a genuine miss.
|
||||
RAW_ITEMS=$(arr_get_cached_items "sonarr" 14400)
|
||||
if [[ -z "$RAW_ITEMS" || "$RAW_ITEMS" == "null" ]]; then
|
||||
info "No fresh cached episode-file data — fetching live (this is the slow path)"
|
||||
SERIES_RESPONSE=$(arr_get_tracked_data "sonarr" "$arr_url" "$arr_key" "v3") || {
|
||||
error "Failed to fetch series from Sonarr — skipping this arr"
|
||||
continue
|
||||
}
|
||||
SERIES_IDS=$(echo "$SERIES_RESPONSE" | jq -r '.[].id')
|
||||
all_tmp=$(mktemp)
|
||||
while IFS= read -r sid; do
|
||||
[[ -z "$sid" ]] && continue
|
||||
_arr_api_append_on_success "$arr_url" "$arr_key" "v3" "episodefile?seriesId=${sid}" "Sonarr" "$all_tmp"
|
||||
done <<< "$SERIES_IDS"
|
||||
RAW_ITEMS=$(jq -s 'add // []' "$all_tmp" 2>/dev/null)
|
||||
rm -f "$all_tmp"
|
||||
arr_item_cache_write "sonarr" "$RAW_ITEMS"
|
||||
fi
|
||||
ITEMS=$(echo "$RAW_ITEMS" | jq -c \
|
||||
'[.[] | {path, file_id:.id, parent_id:.episodeId, title:(.sceneName // .relativePath // .path)}]')
|
||||
else
|
||||
RAW_ITEMS=$(arr_get_cached_items "radarr" 14400)
|
||||
if [[ -z "$RAW_ITEMS" || "$RAW_ITEMS" == "null" ]]; then
|
||||
info "No fresh cached movie-file data — fetching live (this is the slow path)"
|
||||
RAW_ITEMS=$(fetch_radarr_items)
|
||||
arr_item_cache_write "radarr" "$RAW_ITEMS"
|
||||
fi
|
||||
ITEMS=$(echo "$RAW_ITEMS" | jq -c \
|
||||
'[.[] | {path, file_id:.id, parent_id:.movieId, title:(.sceneName // .relativePath // .path)}]')
|
||||
fi
|
||||
|
||||
ITEM_COUNT=$(echo "$ITEMS" | jq 'length' 2>/dev/null)
|
||||
if [[ -z "$ITEM_COUNT" || "$ITEM_COUNT" -eq 0 ]]; then
|
||||
error "0 tracked files for ${arr^} — skipping this arr"
|
||||
continue
|
||||
fi
|
||||
info "$ITEM_COUNT tracked files"
|
||||
|
||||
# ──────────────────────────────────────────────────────────────────────────────────────
|
||||
# Scan
|
||||
# ──────────────────────────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Scanning ━━━"
|
||||
|
||||
SCANNED=0
|
||||
SKIPPED_CACHED=0
|
||||
SKIPPED_UNMAPPED=0
|
||||
CORRUPT_COUNT=0
|
||||
STRIKE_HELD=0
|
||||
REMEDIATED=0
|
||||
REMEDIATE_FAILED=0
|
||||
PROBE_ERRORS=0
|
||||
SUSPECT_COUNT=0
|
||||
CONSECUTIVE_BAD=0
|
||||
QUEUE_PATH=()
|
||||
QUEUE_STRIKES=()
|
||||
QUEUE_ITEM=()
|
||||
|
||||
FRESH_CLEAN_TMP=$(mktemp)
|
||||
|
||||
while IFS= read -r item; do
|
||||
api_path=$(echo "$item" | jq -r '.path')
|
||||
file_id=$(echo "$item" | jq -r '.file_id')
|
||||
parent_id=$(echo "$item" | jq -r '.parent_id')
|
||||
|
||||
host_path=$(translate_path "$api_path")
|
||||
[[ -f "$host_path" ]] || continue
|
||||
[[ -n "$PATH_FILTER" && "$host_path" != *"$PATH_FILTER"* ]] && continue
|
||||
|
||||
stamp="$(stat -c '%Y:%s' "$host_path" 2>/dev/null)"
|
||||
[[ -z "$stamp" ]] && continue
|
||||
|
||||
if [[ "${CLEAN_STATE[$host_path]:-}" == "$stamp" ]]; then
|
||||
(( SKIPPED_CACHED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
(( SCANNED++ ))
|
||||
if [[ "$SCAN_LIMIT" -gt 0 && "$SCANNED" -gt "$SCAN_LIMIT" ]]; then
|
||||
(( SCANNED-- ))
|
||||
break
|
||||
fi
|
||||
|
||||
result=$(probe_file "$host_path")
|
||||
|
||||
if [[ "$result" == "unmapped" ]]; then
|
||||
(( SKIPPED_UNMAPPED++ ))
|
||||
[[ "$ENABLE_LOGGING" == true ]] && warn " ? $host_path — no FFPROBE_PATH_MAP entry covers this share"
|
||||
continue
|
||||
fi
|
||||
|
||||
# A docker-level failure is not evidence about the file. Count it, never queue it.
|
||||
if [[ "$result" == probe_error:* ]]; then
|
||||
(( PROBE_ERRORS++ ))
|
||||
(( CONSECUTIVE_BAD++ ))
|
||||
warn " ? $host_path — probe failed, NOT counted as corrupt: ${result#probe_error:}"
|
||||
if (( CONSECUTIVE_BAD >= CORRUPTION_SCAN_MAX_CONSECUTIVE )); then
|
||||
abort_scan "$CONSECUTIVE_BAD files in a row failed to probe cleanly"
|
||||
fi
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$result" == "clean" ]]; then
|
||||
CONSECUTIVE_BAD=0
|
||||
reset_scan_strikes "$host_path"
|
||||
echo -e "${host_path}\t${stamp}" >> "$FRESH_CLEAN_TMP"
|
||||
[[ "$ENABLE_LOGGING" == true ]] && echo " $ICON_SUCCESS $host_path"
|
||||
continue
|
||||
fi
|
||||
|
||||
# A successful probe that merely warned. Proves the container is alive, so it clears
|
||||
# the consecutive-failure tripwire, but it never becomes a strike.
|
||||
if [[ "$result" == suspect:* ]]; then
|
||||
CONSECUTIVE_BAD=0
|
||||
(( SUSPECT_COUNT++ ))
|
||||
[[ "$ENABLE_LOGGING" == true ]] && warn " ~ $host_path — ffprobe warning (exit 0), NOT corrupt: ${result#suspect:}"
|
||||
continue
|
||||
fi
|
||||
|
||||
# corrupt:<reason> — queued, NOT committed. Nothing reaches the strike file and nothing
|
||||
# is deleted until this arr has been fully probed and the guards below have passed. A
|
||||
# container that dies mid-scan makes every remaining file read as corrupt, and a delete
|
||||
# cannot be undone — so the destructive half has to wait until the corrupt rate for the
|
||||
# whole run is known. 2026-08-23: one Jellyfin restart produced 103 false positives.
|
||||
reason="${result#corrupt:}"
|
||||
(( CORRUPT_COUNT++ ))
|
||||
(( CONSECUTIVE_BAD++ ))
|
||||
|
||||
prev_strikes=$(get_scan_strikes "$host_path")
|
||||
prev_strikes="${prev_strikes//[^0-9]/}"
|
||||
strikes=$(( ${prev_strikes:-0} + 1 ))
|
||||
|
||||
QUEUE_PATH+=("$host_path")
|
||||
QUEUE_STRIKES+=("$strikes")
|
||||
QUEUE_ITEM+=("$item")
|
||||
|
||||
echo " $ICON_ERROR CORRUPT: $host_path (strike $strikes/$CORRUPTION_SCAN_STRIKE_LIMIT)"
|
||||
[[ "$ENABLE_LOGGING" == true ]] && echo " $reason"
|
||||
|
||||
if (( CONSECUTIVE_BAD >= CORRUPTION_SCAN_MAX_CONSECUTIVE )); then
|
||||
abort_scan "$CONSECUTIVE_BAD files in a row failed to probe cleanly"
|
||||
fi
|
||||
done < <(echo "$ITEMS" | jq -c '.[]')
|
||||
|
||||
# ━━━ False-positive guards — run before anything is committed ━━━
|
||||
if (( CORRUPT_COUNT > 0 )); then
|
||||
# The pre-flight check only proves the container was up when the scan started.
|
||||
# Re-check now: a mid-scan death is exactly what this guard exists to catch.
|
||||
check_container_health "$FFPROBE_CONTAINER" "${DOCKER_TIMEOUT:-30}" "Arr Corruption Scan"
|
||||
|
||||
if (( SCANNED >= CORRUPTION_SCAN_GUARD_MIN_SCANNED )); then
|
||||
corrupt_pct=$(( CORRUPT_COUNT * 100 / SCANNED ))
|
||||
if (( corrupt_pct >= CORRUPTION_SCAN_MAX_CORRUPT_PCT )); then
|
||||
abort_scan "$CORRUPT_COUNT of $SCANNED probed files (${corrupt_pct}%) read as corrupt — at or above the ${CORRUPTION_SCAN_MAX_CORRUPT_PCT}% ceiling"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
# ━━━ Guards passed — commit strikes, then remediate whatever reached the limit ━━━
|
||||
for _q in "${!QUEUE_PATH[@]}"; do
|
||||
host_path="${QUEUE_PATH[$_q]}"
|
||||
strikes="${QUEUE_STRIKES[$_q]}"
|
||||
item="${QUEUE_ITEM[$_q]}"
|
||||
|
||||
set_scan_strikes "$host_path" "$strikes"
|
||||
|
||||
[[ "$REMEDIATE" != true ]] && continue
|
||||
|
||||
if (( strikes < CORRUPTION_SCAN_STRIKE_LIMIT )); then
|
||||
warn " $host_path — strike $strikes/$CORRUPTION_SCAN_STRIKE_LIMIT, not yet remediating (needs repeat confirmation)"
|
||||
(( STRIKE_HELD++ ))
|
||||
continue
|
||||
fi
|
||||
reset_scan_strikes "$host_path"
|
||||
|
||||
file_id=$(echo "$item" | jq -r '.file_id')
|
||||
parent_id=$(echo "$item" | jq -r '.parent_id')
|
||||
title=$(echo "$item" | jq -r '.title')
|
||||
|
||||
http_code=$(curl -sf -o /dev/null -w "%{http_code}" -X DELETE \
|
||||
--max-time 15 -H "X-Api-Key: $arr_key" \
|
||||
"${arr_url}/api/v3/${ARR_FILE_ENDPOINT[$arr]}/${file_id}" 2>/dev/null)
|
||||
|
||||
if [[ "$http_code" != "200" ]]; then
|
||||
error " ✗ $title — delete failed (HTTP $http_code)"
|
||||
(( REMEDIATE_FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
sleep 2
|
||||
verify_hasfile=$(arr_api "$arr_url" "$arr_key" "v3" "${ARR_PARENT_ENDPOINT[$arr]}/${parent_id}" "${arr^}" 2>/dev/null \
|
||||
| jq -r '.hasFile // "unknown"')
|
||||
|
||||
if [[ "$verify_hasfile" != "false" ]]; then
|
||||
error " ✗ $title — deleted but hasFile still '$verify_hasfile' — not searching, needs review"
|
||||
(( REMEDIATE_FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
search_code=$(curl -sf -o /dev/null -w "%{http_code}" -X POST \
|
||||
--max-time 30 -H "X-Api-Key: $arr_key" -H "Content-Type: application/json" \
|
||||
-d "{\"name\":\"${ARR_SEARCH_COMMAND[$arr]}\",\"${ARR_SEARCH_ID_FIELD[$arr]}\":[${parent_id}]}" \
|
||||
"${arr_url}/api/v3/command" 2>/dev/null)
|
||||
|
||||
if [[ "$search_code" == "200" || "$search_code" == "201" ]]; then
|
||||
echo " $ICON_SUCCESS $title — deleted, verified, search triggered"
|
||||
(( REMEDIATED++ ))
|
||||
else
|
||||
warn " $title — deleted and verified, but search trigger returned HTTP $search_code"
|
||||
(( REMEDIATE_FAILED++ ))
|
||||
fi
|
||||
done
|
||||
|
||||
merge_clean_state "$FRESH_CLEAN_TMP"
|
||||
rm -f "$FRESH_CLEAN_TMP"
|
||||
|
||||
ARR_SCANNED[$arr]=$SCANNED
|
||||
ARR_SKIPPED_CACHED[$arr]=$SKIPPED_CACHED
|
||||
ARR_SKIPPED_UNMAPPED[$arr]=$SKIPPED_UNMAPPED
|
||||
ARR_CORRUPT[$arr]=$CORRUPT_COUNT
|
||||
ARR_SUSPECT[$arr]=$SUSPECT_COUNT
|
||||
ARR_PROBE_ERRORS[$arr]=$PROBE_ERRORS
|
||||
ARR_STRIKE_HELD[$arr]=$STRIKE_HELD
|
||||
ARR_REMEDIATED[$arr]=$REMEDIATED
|
||||
ARR_REMEDIATE_FAILED[$arr]=$REMEDIATE_FAILED
|
||||
|
||||
(( TOTAL_SCANNED += SCANNED ))
|
||||
(( TOTAL_CORRUPT += CORRUPT_COUNT ))
|
||||
(( TOTAL_REMEDIATED += REMEDIATED ))
|
||||
(( TOTAL_REMEDIATE_FAILED += REMEDIATE_FAILED ))
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY CORRUPTION SCAN SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
for arr in sonarr radarr; do
|
||||
[[ -z "${ARR_SCANNED[$arr]:-}" ]] && continue
|
||||
echo ""
|
||||
echo " ${arr^}:"
|
||||
echo " $ICON_SYNC Newly scanned: ${ARR_SCANNED[$arr]}"
|
||||
echo " $ICON_SUCCESS Skipped (cached): ${ARR_SKIPPED_CACHED[$arr]}"
|
||||
echo " $ICON_WARN Skipped (unmapped): ${ARR_SKIPPED_UNMAPPED[$arr]}"
|
||||
echo " $ICON_ERROR Corrupt found: ${ARR_CORRUPT[$arr]}"
|
||||
echo " $ICON_WARN Warnings (exit 0): ${ARR_SUSPECT[$arr]} (reported, never remediated)"
|
||||
echo " $ICON_WARN Probe errors: ${ARR_PROBE_ERRORS[$arr]} (not counted as corrupt)"
|
||||
if [[ "$REMEDIATE" == true ]]; then
|
||||
echo " $ICON_WARN Held (strikes): ${ARR_STRIKE_HELD[$arr]}"
|
||||
echo " $ICON_SUCCESS Remediated: ${ARR_REMEDIATED[$arr]}"
|
||||
echo " $ICON_ERROR Remediation failed: ${ARR_REMEDIATE_FAILED[$arr]}"
|
||||
fi
|
||||
done
|
||||
echo ""
|
||||
echo " Total:"
|
||||
echo " $ICON_SYNC Newly scanned: $TOTAL_SCANNED"
|
||||
echo " $ICON_ERROR Corrupt found: $TOTAL_CORRUPT"
|
||||
if [[ "$REMEDIATE" == true ]]; then
|
||||
echo " $ICON_SUCCESS Remediated: $TOTAL_REMEDIATED"
|
||||
echo " $ICON_ERROR Remediation failed: $TOTAL_REMEDIATE_FAILED"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
exit 0
|
||||
Executable
+537
@@ -0,0 +1,537 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================== Arr Download Orphan Cleaner =======================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Clears orphaned completed downloads out of the SABnzbd Completed folders that Sonarr and
|
||||
# Radarr import from. Anything sitting there that the arr's queue no longer references is an
|
||||
# orphan — the arr will never touch it again on its own, so without this script the folder
|
||||
# only ever grows.
|
||||
#
|
||||
# Built 2026-07-26 after exactly that: 755G of orphaned completed TV downloads (oldest from
|
||||
# 2022) had silently accumulated and filled the cache pool to 89%. Every existing cleaner
|
||||
# covers the library side (sonarr_cleanup.sh walks SONARR_TV_ROOT etc.) — nothing covered
|
||||
# the download side. This is that missing piece, using the same triage that recovered the
|
||||
# pool that day.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Runs Sonarr then Radarr, sequentially. Per arr, every top-level entry in the configured
|
||||
# download dir is classified:
|
||||
#
|
||||
# TRACKED — basename matches a queue record's outputPath or title → leave it alone,
|
||||
# the arr still knows about it
|
||||
# RECENT — mtime under DOWNLOAD_ORPHAN_AGE days → skip, may be mid-import
|
||||
# JUNK — no video file over DOWNLOAD_ORPHAN_MIN_VIDEO_MB → delete (par2 debris,
|
||||
# samples, failed/never-extracted archives — the arr can't import these,
|
||||
# and if the content is still wanted its own missing-search re-grabs it)
|
||||
# REDUNDANT — parse API matches it AND the library already has every episode / the
|
||||
# movie file → delete (the arr already refused it as not-an-upgrade)
|
||||
# IMPORTABLE — parse API matches it but the library is missing episodes / the movie
|
||||
# → trigger DownloadedEpisodesScan/DownloadedMoviesScan on the folder and
|
||||
# leave it; whatever imports gets swept as REDUNDANT next run, whatever
|
||||
# the arr rejects (XEM-blocked, season-span files) stays HELD for a human
|
||||
# UNMATCHED — parse API can't match it (series/movie not in the arr) → delete. Past the
|
||||
# age gate an entry the arr cannot even name is not going to import: if the
|
||||
# title is in the library and monitored, clearing it lets the arr search a
|
||||
# copy it can actually parse; if it is not in the library, nothing is
|
||||
# tracking it and it is dead weight either way. Guarded — see Non-Empty
|
||||
# Library Requirement below
|
||||
# HELD — IMPORTABLE entries the arr keeps refusing (XEM-blocked, season-span
|
||||
# files), and everything skipped by a guard → report only, human call
|
||||
#
|
||||
# The queue fetch is a hard gate: if it fails, the whole arr is skipped — with no queue
|
||||
# there is no way to tell tracked from orphaned, and guessing means deleting active imports.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Classify Before Acting
|
||||
# Every entry is placed in exactly one class before anything is deleted, and each class
|
||||
# has its own justification. Nothing is removed because it merely looked unwanted — it is
|
||||
# removed because it matched a category whose deletion rationale is written down above.
|
||||
#
|
||||
# The Queue Is the Source of Truth
|
||||
# Tracked-vs-orphaned is decided by the arr's own queue, never inferred from filenames or
|
||||
# timestamps. If the queue cannot be read, the arr is skipped entirely rather than
|
||||
# falling back to a weaker signal — a guess here deletes an active import.
|
||||
#
|
||||
# Deleting Is Recoverable, Deleting Wrong Is Not
|
||||
# The classes that get deleted are ones the arr can re-acquire: junk it could never
|
||||
# import, content the library already has, and entries it cannot even name. Anything
|
||||
# whose loss would be permanent or ambiguous is held and reported for a human instead.
|
||||
#
|
||||
# Abnormal Volume Means Broken Input
|
||||
# The delete cap exists because the realistic failure mode is bad input, not bad logic —
|
||||
# a partial queue fetch classifies live downloads as orphans, and the only visible
|
||||
# symptom is an unusually large delete total. The cap turns that into a stop.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Download folders are written by container users; removing them requires root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock "wait" with an EXIT trap releasing all locks, so an interrupted run never
|
||||
# strands a lock and the daily orchestrator is never silently skipped.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() runs before any HOST*_-prefixed download dir is resolved.
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_CLEANER_ENABLED Toggle
|
||||
# Master switch — exits cleanly when disabled.
|
||||
#
|
||||
# Download Path Restriction
|
||||
# The download dir must exist and live under /mnt/. Anything else is refused rather
|
||||
# than walked, so a blank or malformed path can never point the scan at the filesystem
|
||||
# root or a system directory.
|
||||
#
|
||||
# Queue Fetch Hard Gate
|
||||
# The arr is skipped entirely if its queue cannot be read. Without the queue there is no
|
||||
# way to distinguish tracked from orphaned, and guessing deletes active imports.
|
||||
#
|
||||
# Age Gate
|
||||
# Nothing under DOWNLOAD_ORPHAN_AGE days is touched, so an entry mid-import is never a
|
||||
# deletion candidate regardless of how it classifies.
|
||||
#
|
||||
# Deletion Class Restriction
|
||||
# Only JUNK, parse-verified REDUNDANT and UNMATCHED are deleted. IMPORTABLE entries and
|
||||
# anything a guard has held are reported, never removed.
|
||||
#
|
||||
# Non-Empty Library Requirement
|
||||
# UNMATCHED deletions require the arr to report a non-empty library. An empty or
|
||||
# restoring database answers every parse with "no match", which would condemn the whole
|
||||
# download dir. The library is queried directly rather than inferred from this run's own
|
||||
# match rate — a small batch that is legitimately all-unmatched is normal once daily runs
|
||||
# have caught up, and would otherwise read as a broken database.
|
||||
#
|
||||
# Delete Volume Cap
|
||||
# A run total over DOWNLOAD_ORPHAN_MAX_DELETE_GB aborts the delete pass and notifies.
|
||||
# --i-know-what-im-doing overrides it for a deliberate first run against a known backlog.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run classifies everything and reports, deleting and importing nothing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf (resolved per host after detect_hosts())
|
||||
#
|
||||
# HOST*_SONARR_DOWNLOAD_DIR / HOST*_RADARR_DOWNLOAD_DIR
|
||||
# Host-side path to the arr's completed-download folder. Absent means that arr's
|
||||
# cleanup is skipped, not an error.
|
||||
#
|
||||
# HOST*_SONARR_DOWNLOAD_CONTAINER_DIR / HOST*_RADARR_DOWNLOAD_CONTAINER_DIR
|
||||
# The same folder as the arr container sees it — used when triggering the
|
||||
# DownloadedEpisodesScan / DownloadedMoviesScan path.
|
||||
#
|
||||
# SONARR_URL / SONARR_API_KEY / RADARR_URL / RADARR_API_KEY
|
||||
# Aliased by detect_hosts(). A missing URL or key skips that arr.
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_CLEANER_ENABLED
|
||||
# Master toggle (default: true)
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_AGE
|
||||
# Days before an entry is eligible at all — younger entries may be mid-import
|
||||
# (default: 7)
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_MIN_VIDEO_MB
|
||||
# An entry with no video file above this size is JUNK (default: 50). Sonarr/Radarr only.
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_MIN_AUDIO_MB
|
||||
# The same test for Lidarr (default: 2). Separate because a 50M floor would mark
|
||||
# every album folder as JUNK — single tracks rarely reach it.
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_KEEP_MARKER
|
||||
# A file with this name inside a download folder pins it — the folder is never
|
||||
# classified or deleted (default: .vv-keep). For lossless rips the library holds
|
||||
# only at lower quality, which REDUNDANT would otherwise sweep.
|
||||
#
|
||||
# DOWNLOAD_ORPHAN_MAX_DELETE_GB
|
||||
# Per-run delete budget in GB (default: 100). A backlog above this is drained
|
||||
# safest-first (JUNK, then REDUNDANT, then UNMATCHED) up to the budget, and the
|
||||
# remainder is deferred to the next run rather than aborting the pass.
|
||||
#
|
||||
# SONARR_EXTENSIONS / RADARR_EXTENSIONS
|
||||
# Video extensions used to decide whether an entry contains real media
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arr_download_orphan_cleaner.sh — daily orchestrator entry
|
||||
# arr_download_orphan_cleaner.sh --dry-run — classify and report only
|
||||
# arr_download_orphan_cleaner.sh --status — show config and exit
|
||||
# arr_download_orphan_cleaner.sh --i-know-what-im-doing — bypass MAX_DELETE_GB budget
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_destructive_flags "$@"
|
||||
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
for cmd in curl jq; do
|
||||
if ! command -v "$cmd" >/dev/null 2>&1; then
|
||||
error "$cmd not found — required"
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
|
||||
detect_hosts
|
||||
|
||||
if [[ "${DOWNLOAD_ORPHAN_CLEANER_ENABLED:-false}" != true ]]; then
|
||||
info "DOWNLOAD_ORPHAN_CLEANER_ENABLED=false — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
DOWNLOAD_ORPHAN_AGE="${DOWNLOAD_ORPHAN_AGE:-7}"
|
||||
DOWNLOAD_ORPHAN_MIN_VIDEO_MB="${DOWNLOAD_ORPHAN_MIN_VIDEO_MB:-50}"
|
||||
DOWNLOAD_ORPHAN_MIN_AUDIO_MB="${DOWNLOAD_ORPHAN_MIN_AUDIO_MB:-2}"
|
||||
DOWNLOAD_ORPHAN_KEEP_MARKER="${DOWNLOAD_ORPHAN_KEEP_MARKER:-.vv-keep}"
|
||||
DOWNLOAD_ORPHAN_MAX_DELETE_GB="${DOWNLOAD_ORPHAN_MAX_DELETE_GB:-100}"
|
||||
|
||||
if [[ "${SHOW_STATUS:-false}" == true ]]; then
|
||||
echo "━━━━━ $ICON_SUMMARY DOWNLOAD ORPHAN CLEANER STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Enabled: ${DOWNLOAD_ORPHAN_CLEANER_ENABLED}"
|
||||
echo "$ICON_TIME Age gate: ${DOWNLOAD_ORPHAN_AGE}d"
|
||||
echo "$ICON_DISK Junk threshold: ${DOWNLOAD_ORPHAN_MIN_VIDEO_MB}M video / ${DOWNLOAD_ORPHAN_MIN_AUDIO_MB}M audio"
|
||||
echo "$ICON_SHIELD Delete cap: ${DOWNLOAD_ORPHAN_MAX_DELETE_GB}G"
|
||||
echo "$ICON_SHIELD Keep marker: ${DOWNLOAD_ORPHAN_KEEP_MARKER}"
|
||||
for arr in SONARR RADARR LIDARR; do
|
||||
dir_var="${MY_ID}_${arr}_DOWNLOAD_DIR"
|
||||
echo "$ICON_CLEAN ${arr}: ${!dir_var:-<not configured>}"
|
||||
done
|
||||
exit 0
|
||||
fi
|
||||
|
||||
acquire_lock "wait"
|
||||
trap "_release_all_locks" EXIT
|
||||
|
||||
AGE_CUTOFF=$(( $(date +%s) - DOWNLOAD_ORPHAN_AGE * 86400 ))
|
||||
|
||||
TOTAL_DELETED=0
|
||||
TOTAL_DELETED_MB=0
|
||||
TOTAL_HELD=0
|
||||
TOTAL_DEFERRED=0
|
||||
TOTAL_KEPT=0
|
||||
TOTAL_SCANS=0
|
||||
|
||||
echo "━━━━━ $ICON_CLEAN DOWNLOAD ORPHAN CLEANER ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
[[ "$DRY_RUN" == true ]] && echo "$ICON_SKIP DRY RUN — nothing will be deleted or imported"
|
||||
|
||||
for arr in sonarr radarr lidarr; do
|
||||
api_ver="v3"; [[ "$arr" == "lidarr" ]] && api_ver="v1"
|
||||
url_var="${arr^^}_URL"; key_var="${arr^^}_API_KEY"
|
||||
arr_url="${!url_var:-}"; arr_key="${!key_var:-}"
|
||||
dir_var="${MY_ID}_${arr^^}_DOWNLOAD_DIR"
|
||||
cdir_var="${MY_ID}_${arr^^}_DOWNLOAD_CONTAINER_DIR"
|
||||
dl_dir="${!dir_var:-}"; container_dir="${!cdir_var:-}"
|
||||
|
||||
if [[ -z "$arr_url" || -z "$arr_key" || -z "$dl_dir" ]]; then
|
||||
info "${arr^} download cleanup not configured on $MY_ID — skipping"
|
||||
continue
|
||||
fi
|
||||
if [[ "$dl_dir" != /mnt/* || ! -d "$dl_dir" ]]; then
|
||||
warn "${arr^} download dir invalid or missing: $dl_dir — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC ${arr^} — $dl_dir ━━━"
|
||||
|
||||
ver_var="${arr^^}_VERSION_MAJOR"
|
||||
check_arr_version "$arr_url" "$arr_key" "$api_ver" "${!ver_var}" "${arr^}" || {
|
||||
warn "${arr^} version check failed — skipping this arr"
|
||||
continue
|
||||
}
|
||||
# min_mb is per-arr because the JUNK test is "contains no real media file". A 50MB floor
|
||||
# is right for video and catastrophic for audio — most single tracks never reach it, so
|
||||
# every music folder would classify as JUNK and be deleted regardless of import state.
|
||||
case "$arr" in
|
||||
sonarr)
|
||||
queue_endpoint="queue?pageSize=1000&includeUnknownSeriesItems=true"
|
||||
exts_var="SONARR_EXTENSIONS"
|
||||
scan_command="DownloadedEpisodesScan"
|
||||
library_endpoint="series"
|
||||
min_mb="$DOWNLOAD_ORPHAN_MIN_VIDEO_MB"
|
||||
;;
|
||||
radarr)
|
||||
queue_endpoint="queue?pageSize=1000&includeUnknownMovieItems=true"
|
||||
exts_var="RADARR_EXTENSIONS"
|
||||
scan_command="DownloadedMoviesScan"
|
||||
library_endpoint="movie"
|
||||
min_mb="$DOWNLOAD_ORPHAN_MIN_VIDEO_MB"
|
||||
;;
|
||||
lidarr)
|
||||
queue_endpoint="queue?pageSize=1000&includeUnknownArtistItems=true"
|
||||
exts_var="LIDARR_EXTENSIONS"
|
||||
scan_command="DownloadedAlbumsScan"
|
||||
library_endpoint="artist"
|
||||
min_mb="$DOWNLOAD_ORPHAN_MIN_AUDIO_MB"
|
||||
;;
|
||||
esac
|
||||
|
||||
QUEUE_JSON=$(arr_api "$arr_url" "$arr_key" "$api_ver" "$queue_endpoint" "${arr^}") || {
|
||||
error "${arr^} queue fetch failed — cannot tell tracked from orphaned, skipping this arr"
|
||||
continue
|
||||
}
|
||||
|
||||
declare -A PROTECTED=()
|
||||
while IFS= read -r name; do
|
||||
[[ -n "$name" ]] && PROTECTED["$name"]=1
|
||||
done < <(echo "$QUEUE_JSON" | jq -r '.records[] | (.outputPath // empty | split("/") | last), (.title // empty)')
|
||||
|
||||
eval "arr_exts=(\"\${${exts_var}[@]}\")"
|
||||
|
||||
DELETE_PATHS=()
|
||||
DELETE_SIZES=()
|
||||
DELETE_LABELS=()
|
||||
SCAN_PATHS=()
|
||||
UNMATCHED_PATHS=()
|
||||
UNMATCHED_SIZES=()
|
||||
arr_tracked=0; arr_recent=0; arr_held=0; arr_delete_mb=0; arr_kept=0
|
||||
|
||||
while IFS= read -r entry; do
|
||||
base="${entry##*/}"
|
||||
|
||||
# An operator keep-marker outranks every verdict below. Needed because REDUNDANT only
|
||||
# asks "does the library hold this album", not "at what quality" — a lossless rip whose
|
||||
# library copy is MP3 is redundant by that test and would be swept on the next run.
|
||||
# The marker is a file inside the folder rather than a conf list so it survives renames
|
||||
# and cannot drift out of sync with what is actually on disk.
|
||||
if [[ -e "$entry/$DOWNLOAD_ORPHAN_KEEP_MARKER" ]]; then
|
||||
arr_kept=$((arr_kept + 1))
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ -n "${PROTECTED[$base]:-}" ]]; then
|
||||
arr_tracked=$((arr_tracked + 1))
|
||||
continue
|
||||
fi
|
||||
|
||||
mtime=$(stat -c %Y "$entry" 2>/dev/null) || continue
|
||||
if (( mtime > AGE_CUTOFF )); then
|
||||
arr_recent=$((arr_recent + 1))
|
||||
continue
|
||||
fi
|
||||
|
||||
# JUNK means "holds no real media". That verdict is only as good as the extension
|
||||
# list, and a missing extension turns real content into a delete — 2026-08-21 the
|
||||
# audio list had no "wv", which classified 23 folders of WavPack lossless (1.5G per
|
||||
# file) as junk. So a folder with large files that are merely *unrecognised* is held
|
||||
# for review, never deleted; only a folder with nothing big in it at all is junk.
|
||||
has_media=false
|
||||
big_unknown=0
|
||||
while IFS= read -r f; do
|
||||
if has_extension "$f" "${arr_exts[@]}"; then
|
||||
has_media=true
|
||||
break
|
||||
fi
|
||||
big_unknown=$((big_unknown + 1))
|
||||
done < <(find "$entry" -type f -size +"${min_mb}"M 2>/dev/null)
|
||||
|
||||
size_mb=$(dir_size_mb "$entry") || size_mb=0
|
||||
|
||||
if [[ "$has_media" == false ]] && (( big_unknown > 0 )); then
|
||||
warn " no recognised media, but $big_unknown large file(s) of unknown type — holding: $base"
|
||||
arr_held=$((arr_held + 1))
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$has_media" == false ]]; then
|
||||
DELETE_PATHS+=("$entry")
|
||||
DELETE_SIZES+=("$size_mb")
|
||||
DELETE_LABELS+=("JUNK")
|
||||
arr_delete_mb=$((arr_delete_mb + size_mb))
|
||||
continue
|
||||
fi
|
||||
|
||||
enc_title=$(jq -rn --arg t "$base" '$t|@uri')
|
||||
parse=$(arr_api "$arr_url" "$arr_key" "$api_ver" "parse?title=${enc_title}" "${arr^}") || {
|
||||
warn " parse failed for: $base — holding"
|
||||
arr_held=$((arr_held + 1))
|
||||
continue
|
||||
}
|
||||
|
||||
case "$arr" in
|
||||
sonarr)
|
||||
matched=$(echo "$parse" | jq '(.series != null) and ((.episodes | length) > 0)')
|
||||
missing=$(echo "$parse" | jq '[.episodes[]? | select(.hasFile == false)] | length')
|
||||
;;
|
||||
radarr)
|
||||
# Radarr's parse never populates hasFile — movieFileId is the reliable signal
|
||||
matched=$(echo "$parse" | jq '.movie != null')
|
||||
missing=$(echo "$parse" | jq 'if (.movie.movieFileId // 0) > 0 then 0 else 1 end')
|
||||
;;
|
||||
lidarr)
|
||||
# Lidarr's parse returns albums with statistics:null, so the track count has
|
||||
# to be read back from album/{id} — the same shape of gap as Radarr's hasFile.
|
||||
matched=$(echo "$parse" | jq '(.artist != null) and ((.albums | length) > 0)')
|
||||
missing=1
|
||||
if [[ "$matched" == true ]]; then
|
||||
album_id=$(echo "$parse" | jq -r '.albums[0].id // empty')
|
||||
if [[ -z "$album_id" ]]; then
|
||||
warn " parse matched but returned no album id: $base — holding"
|
||||
arr_held=$((arr_held + 1))
|
||||
continue
|
||||
fi
|
||||
album_json=$(arr_api "$arr_url" "$arr_key" "$api_ver" "album/$album_id" "${arr^}") || {
|
||||
warn " album lookup failed for: $base — holding"
|
||||
arr_held=$((arr_held + 1))
|
||||
continue
|
||||
}
|
||||
missing=$(echo "$album_json" | jq 'if ((.statistics.trackFileCount // 0) > 0) then 0 else 1 end')
|
||||
fi
|
||||
;;
|
||||
esac
|
||||
|
||||
if [[ "$matched" != true ]]; then
|
||||
UNMATCHED_PATHS+=("$entry")
|
||||
UNMATCHED_SIZES+=("$size_mb")
|
||||
elif (( missing == 0 )); then
|
||||
DELETE_PATHS+=("$entry")
|
||||
DELETE_SIZES+=("$size_mb")
|
||||
DELETE_LABELS+=("REDUNDANT")
|
||||
arr_delete_mb=$((arr_delete_mb + size_mb))
|
||||
else
|
||||
SCAN_PATHS+=("$base")
|
||||
arr_held=$((arr_held + 1))
|
||||
fi
|
||||
done < <(find "$dl_dir" -mindepth 1 -maxdepth 1 2>/dev/null)
|
||||
|
||||
# An arr with an empty or still-restoring database answers every parse with "no match",
|
||||
# which would turn the whole download dir into UNMATCHED and delete it. Confirm the
|
||||
# library actually holds titles before trusting a no-match to mean what it says. This
|
||||
# has to be asked of the arr directly — inferring health from the run's own matches
|
||||
# fails on a small batch that is legitimately all-unmatched, which is the normal case
|
||||
# once daily runs have caught up.
|
||||
if (( ${#UNMATCHED_PATHS[@]} > 0 )); then
|
||||
library_count=$(arr_api "$arr_url" "$arr_key" "$api_ver" "$library_endpoint" "${arr^}" | jq 'length' 2>/dev/null)
|
||||
if [[ ! "$library_count" =~ ^[0-9]+$ ]] || (( library_count == 0 )); then
|
||||
warn " ${arr^}: library reports ${library_count:-no} titles — cannot trust 'no match', holding ${#UNMATCHED_PATHS[@]} unmatched"
|
||||
arr_held=$((arr_held + ${#UNMATCHED_PATHS[@]}))
|
||||
else
|
||||
for i in "${!UNMATCHED_PATHS[@]}"; do
|
||||
DELETE_PATHS+=("${UNMATCHED_PATHS[$i]}")
|
||||
DELETE_SIZES+=("${UNMATCHED_SIZES[$i]}")
|
||||
DELETE_LABELS+=("UNMATCHED")
|
||||
arr_delete_mb=$((arr_delete_mb + UNMATCHED_SIZES[i]))
|
||||
done
|
||||
fi
|
||||
fi
|
||||
|
||||
# The cap is a per-run risk budget, not a reason to do nothing. Aborting the whole pass
|
||||
# once the backlog exceeds it is self-defeating: the backlog can never shrink below the
|
||||
# cap on its own, so every later run aborts too and the pool fills anyway (exactly how
|
||||
# 347G accumulated here by 2026-08-21). Delete in ascending order of risk instead, stop
|
||||
# at the cap, and defer the rest to the next run so a backlog drains over days.
|
||||
#
|
||||
# Live downloads are already protected by DOWNLOAD_ORPHAN_AGE, not by this cap — anything
|
||||
# in flight is younger than the age gate and never reaches classification. That is what
|
||||
# makes draining safe: the partial-queue-data case the cap was written for cannot put a
|
||||
# still-downloading entry in these arrays.
|
||||
cap_mb=$((DOWNLOAD_ORPHAN_MAX_DELETE_GB * 1024))
|
||||
cap_active=true
|
||||
[[ "$I_KNOW" == true || "$DRY_RUN" == true ]] && cap_active=false
|
||||
|
||||
arr_deferred=0; arr_deferred_mb=0; arr_run_mb=0
|
||||
|
||||
if [[ "$cap_active" == true ]] && (( arr_delete_mb > cap_mb )); then
|
||||
warn " ${arr^}: $((arr_delete_mb / 1024))G classified vs ${DOWNLOAD_ORPHAN_MAX_DELETE_GB}G cap — deleting safest-first up to the cap, deferring the rest"
|
||||
notify "${arr^} download orphan backlog is $((arr_delete_mb / 1024))G on $(hostname), above the ${DOWNLOAD_ORPHAN_MAX_DELETE_GB}G per-run cap. Draining safest-first; the remainder follows on later runs. Re-run with --i-know-what-im-doing to clear it in one pass." \
|
||||
"Download Orphan Cleaner" "warning"
|
||||
fi
|
||||
|
||||
# JUNK first (no media at all), then REDUNDANT (parse-verified already in the library),
|
||||
# then UNMATCHED last — it rests on "the arr does not know this title", the weakest of
|
||||
# the three signals, so it is the first thing the cap defers.
|
||||
for pass in JUNK REDUNDANT UNMATCHED; do
|
||||
for i in "${!DELETE_PATHS[@]}"; do
|
||||
[[ "${DELETE_LABELS[$i]}" == "$pass" ]] || continue
|
||||
entry="${DELETE_PATHS[$i]}"
|
||||
|
||||
if [[ "$cap_active" == true ]] && (( arr_run_mb + DELETE_SIZES[i] > cap_mb )); then
|
||||
arr_deferred=$((arr_deferred + 1))
|
||||
arr_deferred_mb=$((arr_deferred_mb + DELETE_SIZES[i]))
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo " $ICON_SKIP would delete [${DELETE_LABELS[$i]}]: ${entry##*/} (${DELETE_SIZES[$i]}M)"
|
||||
else
|
||||
echo " $ICON_TRASH deleting [${DELETE_LABELS[$i]}]: ${entry##*/} (${DELETE_SIZES[$i]}M)"
|
||||
rm -rf "$entry"
|
||||
fi
|
||||
arr_run_mb=$((arr_run_mb + DELETE_SIZES[i]))
|
||||
TOTAL_DELETED=$((TOTAL_DELETED + 1))
|
||||
TOTAL_DELETED_MB=$((TOTAL_DELETED_MB + DELETE_SIZES[i]))
|
||||
done
|
||||
done
|
||||
|
||||
if (( arr_deferred > 0 )); then
|
||||
echo " $ICON_WARN ${arr^}: deferred $arr_deferred entries ($((arr_deferred_mb / 1024))G) to the next run — cap reached"
|
||||
TOTAL_DEFERRED=$((TOTAL_DEFERRED + arr_deferred))
|
||||
fi
|
||||
|
||||
for base in "${SCAN_PATHS[@]}"; do
|
||||
if [[ -z "$container_dir" ]]; then
|
||||
echo " $ICON_WARN IMPORTABLE but ${cdir_var} not set — holding: $base"
|
||||
continue
|
||||
fi
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo " $ICON_SKIP would trigger $scan_command: $base"
|
||||
else
|
||||
echo " $ICON_RUN triggering $scan_command: $base"
|
||||
payload=$(jq -nc --arg n "$scan_command" --arg p "$container_dir/$base" '{name: $n, path: $p, importMode: "Move"}')
|
||||
curl -sf --max-time 30 -X POST \
|
||||
-H "X-Api-Key: $arr_key" -H "Content-Type: application/json" \
|
||||
-d "$payload" "$arr_url/api/v3/command" >/dev/null \
|
||||
|| warn " $scan_command trigger failed for: $base"
|
||||
TOTAL_SCANS=$((TOTAL_SCANS + 1))
|
||||
fi
|
||||
done
|
||||
|
||||
echo " $ICON_SUMMARY ${arr^}: $arr_tracked tracked, $arr_kept kept, $arr_recent recent, $((${#DELETE_PATHS[@]} - arr_deferred)) deleted ($((arr_run_mb / 1024))G), $arr_deferred deferred ($((arr_deferred_mb / 1024))G), ${#SCAN_PATHS[@]} import scans, $arr_held held"
|
||||
TOTAL_HELD=$((TOTAL_HELD + arr_held))
|
||||
TOTAL_KEPT=$((TOTAL_KEPT + arr_kept))
|
||||
unset PROTECTED
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_DONE SUMMARY ━━━━━"
|
||||
echo "$ICON_TRASH Deleted: $TOTAL_DELETED ($((TOTAL_DELETED_MB / 1024))G)"
|
||||
echo "$ICON_RUN Import scans: $TOTAL_SCANS"
|
||||
echo "$ICON_WARN Held: $TOTAL_HELD"
|
||||
echo "$ICON_SKIP Deferred: $TOTAL_DEFERRED"
|
||||
echo "$ICON_SHIELD Kept (marker): $TOTAL_KEPT"
|
||||
|
||||
# Held alone never notifies — there is always something awaiting review, and on a daily
|
||||
# schedule that would be a notification every morning saying nothing happened.
|
||||
if [[ "$DRY_RUN" != true ]] && (( TOTAL_DELETED > 0 )); then
|
||||
notify "Download orphan cleaner on $(hostname): deleted $TOTAL_DELETED orphans ($((TOTAL_DELETED_MB / 1024))G), triggered $TOTAL_SCANS import scans, $TOTAL_HELD held for review" \
|
||||
"Download Orphan Cleaner" "normal"
|
||||
fi
|
||||
Executable
+246
@@ -0,0 +1,246 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ======================= Arr Full Library Rescan =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Forces a genuine full disk↔database reconciliation for Lidarr/Sonarr/Radarr. Organic
|
||||
# scans (triggered by new imports, RSS sync, etc.) only touch the files actually involved —
|
||||
# an artist/series/movie that already has files sitting untouched on disk never gets its
|
||||
# file-tracking stats refreshed on its own. Confirmed 2026-07-16: Lidarr reported only ~23%
|
||||
# of its true trackFileCount with no active scan running, for 1,004 of 1,357 artists — files
|
||||
# verified present and readable on disk the whole time. Every downstream script (cleanup,
|
||||
# duplicate-artist detection, discovery) trusts these arr stats as source of truth for what's
|
||||
# on the share, so silent drift like this is exactly what check_tracked_count_floor() exists
|
||||
# to catch reactively. This job exists to catch it proactively instead of waiting for someone
|
||||
# to notice a suspiciously low number.
|
||||
#
|
||||
# Runs sequentially across all three arrs, never parallel — each is a heavy full-disk walk,
|
||||
# and running them concurrently would just contend for the same disk I/O for no benefit.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Lidarr, then Sonarr, then Radarr — strictly sequential. Per arr:
|
||||
#
|
||||
# 1. Reachability
|
||||
# → check_api; unreachable skips this arr only
|
||||
#
|
||||
# 2. Already-scanning check
|
||||
# → a rescan already active (manual, or another script) means skip rather than
|
||||
# stack a second full-disk walk on top of it
|
||||
#
|
||||
# 3. Capture the before count
|
||||
# → tracked file count read from the arr's own stats
|
||||
#
|
||||
# 4. Trigger the rescan command
|
||||
# → RescanFolders (Lidarr) / RescanSeries (Sonarr) / RescanMovie (Radarr)
|
||||
#
|
||||
# 5. Poll to completion
|
||||
# → bounded by ARR_FULL_RESCAN_TIMEOUT
|
||||
#
|
||||
# 6. Report the delta
|
||||
# → before vs after tracked count, so drift that was corrected is visible
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Proactive, Not Reactive
|
||||
# check_tracked_count_floor() catches stat drift reactively, at the moment some other
|
||||
# script is about to act on bad numbers. This job exists so that drift is corrected on a
|
||||
# schedule instead of being discovered by whichever cleanup happens to trip over it first.
|
||||
#
|
||||
# Sequential by Design
|
||||
# Each rescan is a full-disk walk. Running three concurrently contends for the same
|
||||
# spindles and finishes no sooner, so the arrs are never parallelised — the slowness is
|
||||
# accepted deliberately rather than optimised into I/O thrash.
|
||||
#
|
||||
# Never Stack a Scan
|
||||
# An already-running rescan is left alone rather than duplicated. A second concurrent
|
||||
# walk of the same library doubles the I/O cost and returns nothing the first will not.
|
||||
#
|
||||
# Per-Arr Isolation
|
||||
# One arr being down, slow, or already scanning must never prevent the other two from
|
||||
# being reconciled. Partial coverage beats a skipped run.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Kept for consistency across Arrs_Stack/ — this script makes no direct filesystem writes.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock "wait" — waits for a prior run rather than skipping or colliding. A full
|
||||
# rescan across three arrs runs long and is worth queuing behind, not silently dropping.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases each arr's URL and API key.
|
||||
#
|
||||
# jq Dependency Check
|
||||
# Fails fast if jq is missing. Both the before/after tracked counts and the command
|
||||
# payload are built with jq — without it the counts read empty and every delta would be
|
||||
# reported as if nothing changed.
|
||||
#
|
||||
# Reachability Check
|
||||
# check_api before touching an arr; unreachable skips that arr only.
|
||||
#
|
||||
# Active-Rescan Check
|
||||
# Skips triggering a new rescan if one is already active on that arr, so a duplicate
|
||||
# full-disk walk is never stacked. See Tools/arr_rescan_monitor.sh for catching that
|
||||
# arr's cache up once the pre-existing scan finishes, rather than waiting a week.
|
||||
#
|
||||
# Sequential Only
|
||||
# Two arrs' rescans never run in parallel.
|
||||
#
|
||||
# Per-Arr Isolation
|
||||
# One arr failing, timing out, or being skipped never blocks the others.
|
||||
#
|
||||
# Timeout Bound
|
||||
# ARR_FULL_RESCAN_TIMEOUT caps the wait per arr, so a rescan that never completes cannot
|
||||
# hold the weekly window open indefinitely.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run reports which arrs would be rescanned and triggers nothing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
# ARR_FULL_RESCAN_TIMEOUT — seconds to wait per arr (default 3600). A whole-library
|
||||
# RescanFolders/RescanSeries/RescanMovie is far heavier than the 600s pre-flight scan
|
||||
# timeout used elsewhere — that shorter timeout is sized for a single release, not a
|
||||
# full-library walk.
|
||||
#
|
||||
# host*.conf
|
||||
# HOST1_LIDARR_URL / _API_KEY, HOST1_SONARR_URL / _API_KEY, HOST1_RADARR_URL / _API_KEY
|
||||
# — aliased by detect_hosts()
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arr_full_rescan.sh — normal run
|
||||
# arr_full_rescan.sh --dry-run — preview which arrs would be rescanned, trigger nothing
|
||||
# arr_full_rescan.sh --log — verbose, per-arr detail
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Both the tracked-count reads and the command payload are built with jq — without it the
|
||||
# counts read empty and every arr would report a zero delta as if nothing had drifted.
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
notify "Arr full rescan failed on $(hostname) — jq not installed" "Arr Full Rescan" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock "wait"
|
||||
|
||||
detect_hosts
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no rescans will be triggered"
|
||||
|
||||
declare -A ARR_FULL_RESCAN_COMMAND=(
|
||||
[lidarr]="RescanFolders"
|
||||
[sonarr]="RescanSeries"
|
||||
[radarr]="RescanMovie"
|
||||
)
|
||||
|
||||
RESCANNED=0
|
||||
SKIPPED=0
|
||||
|
||||
for arr in lidarr sonarr radarr; do
|
||||
url_var="${arr^^}_URL"; key_var="${arr^^}_API_KEY"
|
||||
url="${!url_var:-}"; key="${!key_var:-}"
|
||||
ver="v3"; [[ "$arr" == "lidarr" ]] && ver="v1"
|
||||
|
||||
if [[ -z "$url" || -z "$key" ]]; then
|
||||
info "${arr^} not configured on $MY_ID — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
check_api "$url" "${arr^}" 10 || {
|
||||
warn "${arr^} unreachable — skipping full rescan this run"
|
||||
(( SKIPPED++ ))
|
||||
continue
|
||||
}
|
||||
|
||||
active=$(arr_active_rescan_command "$arr" "$url" "$key" "$ver")
|
||||
if [[ -n "$active" ]]; then
|
||||
warn "${arr^} already mid-rescan ($active) — skipping, will catch it next scheduled run (run Tools/arr_rescan_monitor.sh ${arr} to refresh its cache as soon as this one finishes instead of waiting)"
|
||||
(( SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
endpoint="${ARR_LIBRARY_ENDPOINT[$arr]}"
|
||||
expr="${ARR_TRACKED_COUNT_EXPR[$arr]}"
|
||||
|
||||
before_json=$(curl -sf --max-time 60 -H "X-Api-Key: $key" "${url}/api/${ver}/${endpoint}" 2>/dev/null)
|
||||
before_count=$(echo "$before_json" | jq "$expr" 2>/dev/null)
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would trigger full ${ARR_FULL_RESCAN_COMMAND[$arr]} for ${arr} (currently: ${before_count:-unknown} tracked)"
|
||||
continue
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Triggering full ${ARR_FULL_RESCAN_COMMAND[$arr]} for ${arr} (before: ${before_count:-unknown} tracked)"
|
||||
payload=$(jq -c -n --arg name "${ARR_FULL_RESCAN_COMMAND[$arr]}" '{name:$name}')
|
||||
trigger_and_await_command "$url" "$key" "$ver" "$payload" "${ARR_FULL_RESCAN_TIMEOUT:-3600}" "$arr"
|
||||
|
||||
after_json=$(curl -sf --max-time 60 -H "X-Api-Key: $key" "${url}/api/${ver}/${endpoint}" 2>/dev/null)
|
||||
after_count=$(echo "$after_json" | jq "$expr" 2>/dev/null)
|
||||
|
||||
if [[ -n "$after_json" && -n "$after_count" && "$after_count" != "null" ]]; then
|
||||
arr_cache_write "$arr" "$after_json"
|
||||
log "$ICON_DONE ${arr^} rescan complete — tracked: ${before_count:-?} → ${after_count}"
|
||||
(( RESCANNED++ ))
|
||||
|
||||
# A completed full rescan is ground truth — if it's STILL far below the running
|
||||
# baseline, that's a real problem (missing disk, permissions, actual data loss),
|
||||
# not a stale-cache or mid-scan artifact. Worth a direct heads-up either way.
|
||||
count_file_var="${arr^^}_TRACKED_COUNT_FILE"
|
||||
min_pct_var="${arr^^}_MIN_TRACKED_PCT"
|
||||
count_file="${!count_file_var:-}"
|
||||
min_pct="${!min_pct_var:-80}"
|
||||
if [[ -n "$count_file" && -f "$count_file" ]]; then
|
||||
baseline=$(cat "$count_file" 2>/dev/null || echo 0)
|
||||
if [[ "$baseline" -gt 0 ]]; then
|
||||
pct=$(awk "BEGIN {printf \"%d\", ($after_count / $baseline) * 100}")
|
||||
if [[ "$pct" -lt "$min_pct" ]]; then
|
||||
notify "${arr^} full rescan complete but tracked count still ${pct}% of baseline ($after_count vs $baseline) on $(hostname) — real drop, not a scan artifact, needs a look" \
|
||||
"Arr Full Rescan" "warning"
|
||||
else
|
||||
echo "$after_count" > "$count_file"
|
||||
fi
|
||||
else
|
||||
echo "$after_count" > "$count_file"
|
||||
fi
|
||||
fi
|
||||
else
|
||||
warn "${arr^} rescan finished but re-fetch failed — cache not updated"
|
||||
fi
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY ARR FULL RESCAN SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Rescanned: $RESCANNED"
|
||||
echo "$ICON_SKIP Skipped: $SKIPPED"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
exit 0
|
||||
@@ -10,6 +10,12 @@
|
||||
# rsync in the weekly sync window: once arrs agree on what to track, rsync
|
||||
# spreads the actual files.
|
||||
#
|
||||
# The LOCAL side of each sync (_local_library()) is cache-first (2026-07-17) — comes from
|
||||
# the shared tracked-data cache via arr_get_tracked_data(), fresh (kept warm every 30min by
|
||||
# arr_cache_prefill.sh), live fetch as fallback. The REMOTE side (_remote_library()) is
|
||||
# unaffected — that cache is per-host by design, so a remote node's library is always
|
||||
# fetched live here.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
@@ -37,11 +43,13 @@
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Remote API Keys Never Stored
|
||||
# SSHes to each remote node and reads the key directly from that arr's
|
||||
# config.xml in its appdata directory. Only the API response (JSON) is
|
||||
# returned — the key never leaves the remote node. Self-maintaining: key
|
||||
# regeneration on the remote is picked up automatically next run.
|
||||
# Remote API Access — Cache-First, SSH Fallback
|
||||
# If conf_sync.sh has populated /tmp/varaverk/conf/ and
|
||||
# load_config.sh has sourced it, HOST*_<ARR>_API_KEY vars are available
|
||||
# in the environment. Remote functions use them to call the arr API
|
||||
# directly over Tailscale (no SSH, no remote shell). If the cached key
|
||||
# is absent (first boot, cache not yet populated) the functions fall back
|
||||
# to SSHing in and reading the key from config.xml on the remote node.
|
||||
#
|
||||
# Blocklist TSV
|
||||
# ARR_SYNC_BLOCKLIST in DATA_DIR tombstones IDs that must never be re-added
|
||||
@@ -122,6 +130,7 @@ BLOCKLIST_ACTION=""
|
||||
BLOCKLIST_ARR=""
|
||||
BLOCKLIST_ID=""
|
||||
BLOCKLIST_REASON=""
|
||||
FORCE_MODE=false
|
||||
FILTERED_ARGS=()
|
||||
_skip_next=false
|
||||
for _arg in "$@"; do
|
||||
@@ -130,6 +139,7 @@ for _arg in "$@"; do
|
||||
--blocklist-add) BLOCKLIST_ACTION="add" ;;
|
||||
--blocklist-remove) BLOCKLIST_ACTION="remove" ;;
|
||||
--blocklist-list) BLOCKLIST_ACTION="list" ;;
|
||||
--force) FORCE_MODE=true ;;
|
||||
*) FILTERED_ARGS+=("$_arg") ;;
|
||||
esac
|
||||
done
|
||||
@@ -170,13 +180,14 @@ done
|
||||
unset _tool
|
||||
|
||||
detect_hosts
|
||||
require_partnership
|
||||
|
||||
# ── Runtime config with defaults ──────────────────────────────────────────────────────────────
|
||||
ARR_SYNC_ENABLED="${ARR_SYNC_ENABLED:-true}"
|
||||
ARR_SYNC_BLOCKLIST="${ARR_SYNC_BLOCKLIST:-${DATA_DIR}/arr_sync_blocklist.tsv}"
|
||||
ARR_SYNC_CONNECT_TIMEOUT="${ARR_SYNC_CONNECT_TIMEOUT:-10}"
|
||||
ARR_SYNC_API_TIMEOUT="${ARR_SYNC_API_TIMEOUT:-60}"
|
||||
DOCKER_APPDATA_BASE="${DOCKER_APPDATA_BASE:-/mnt/user/appdata}"
|
||||
|
||||
ARR_SYNC_LIDARR_PORT="${ARR_SYNC_LIDARR_PORT:-8686}"
|
||||
ARR_SYNC_SONARR_PORT="${ARR_SYNC_SONARR_PORT:-8989}"
|
||||
ARR_SYNC_RADARR_PORT="${ARR_SYNC_RADARR_PORT:-7878}"
|
||||
@@ -186,6 +197,9 @@ if [[ "$ARR_SYNC_ENABLED" != "true" ]]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Config: api-timeout=${ARR_SYNC_API_TIMEOUT}s connect-timeout=${ARR_SYNC_CONNECT_TIMEOUT}s blocklist=${ARR_SYNC_BLOCKLIST}"
|
||||
log "$ICON_GEAR Ports: lidarr=${ARR_SYNC_LIDARR_PORT} sonarr=${ARR_SYNC_SONARR_PORT} radarr=${ARR_SYNC_RADARR_PORT}"
|
||||
|
||||
# ── Arr type definitions ───────────────────────────────────────────────────────────────────────
|
||||
# Each arr type maps to its port, API version, endpoint, stable ID field, and display name field
|
||||
declare -A _PORT=([lidarr]="$ARR_SYNC_LIDARR_PORT" [sonarr]="$ARR_SYNC_SONARR_PORT" [radarr]="$ARR_SYNC_RADARR_PORT")
|
||||
@@ -198,13 +212,7 @@ declare -A _ID_TYPE=([lidarr]="string" [sonarr]="int" [radarr]="int")
|
||||
ARR_TYPES=(lidarr sonarr radarr)
|
||||
|
||||
# ── Remote node discovery ──────────────────────────────────────────────────────────────────────
|
||||
REMOTE_NODES=()
|
||||
for _hv in HOST1 HOST2 HOST3 HOST4 HOST5 HOST6 HOST7 HOST8; do
|
||||
[[ "$_hv" == "$MY_ID" ]] && continue
|
||||
[[ -z "${!_hv:-}" ]] && continue
|
||||
REMOTE_NODES+=("$_hv")
|
||||
done
|
||||
unset _hv
|
||||
discover_remote_nodes
|
||||
|
||||
if [[ "${#REMOTE_NODES[@]}" -eq 0 ]]; then
|
||||
warn "No remote nodes defined in master.conf — nothing to sync"
|
||||
@@ -308,19 +316,16 @@ _delete_local_item() {
|
||||
"${url}/api/${api_ver}/${endpoint}/${internal_id}?deleteFiles=false" 2>/dev/null
|
||||
}
|
||||
|
||||
# Delete item from remote arr by stable_id via SSH.
|
||||
# Outputs: HTTP code on success | "not_found" if item absent | empty on SSH/API failure.
|
||||
# Delete item from remote arr by stable_id.
|
||||
# Outputs: HTTP code on success | "not_found" if item absent | empty on failure.
|
||||
# deleteFiles=false — files become orphans for arr_cleanup to handle with its safety checks.
|
||||
# Args: node_id port api_ver arr_type endpoint id_field id_type stable_id
|
||||
_delete_remote_item() {
|
||||
local node_id="$1" port="$2" api_ver="$3" endpoint="$4" config_xml="$5"
|
||||
local node_id="$1" port="$2" api_ver="$3" arr_type="$4" endpoint="$5"
|
||||
local id_field="$6" id_type="$7" stable_id="$8"
|
||||
local node_name="${!node_id}"
|
||||
local ts_name="${node_name,,}"
|
||||
local node_ip
|
||||
node_ip=$(tailscale ip -4 "$ts_name" 2>/dev/null)
|
||||
[[ -z "$node_ip" ]] && node_ip=$(tailscale status 2>/dev/null | \
|
||||
awk -v n="$ts_name" '$2 ~ "^" n { print $1; exit }')
|
||||
[[ -z "$node_ip" ]] && return 1
|
||||
node_ip=$(resolve_tailscale_ip "$node_name") || return 1
|
||||
|
||||
local select_expr
|
||||
if [[ "$id_type" == "string" ]]; then
|
||||
@@ -329,6 +334,21 @@ _delete_remote_item() {
|
||||
select_expr=".[] | select(.${id_field} == ${stable_id}) | .id"
|
||||
fi
|
||||
|
||||
local _kvar="${node_id}_${arr_type^^}_API_KEY"; local cached_key="${!_kvar:-}"
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
local library internal_id
|
||||
library=$(curl -sf --max-time "$ARR_SYNC_API_TIMEOUT" \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
"http://${node_ip}:${port}/api/${api_ver}/${endpoint}" 2>/dev/null)
|
||||
[[ -z "$library" ]] && return 1
|
||||
internal_id=$(echo "$library" | jq -r "${select_expr}" 2>/dev/null | head -1)
|
||||
[[ -z "$internal_id" ]] && echo "not_found" && return 0
|
||||
curl -sf -o /dev/null -w '%{http_code}' -X DELETE \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
"http://${node_ip}:${port}/api/${api_ver}/${endpoint}/${internal_id}?deleteFiles=false" 2>/dev/null
|
||||
return
|
||||
fi
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/${arr_type^}/config.xml"
|
||||
ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" bash <<REMOTE 2>/dev/null
|
||||
KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
@@ -382,7 +402,6 @@ if [[ -n "$BLOCKLIST_ACTION" ]]; then
|
||||
_bl_id_field="${_ID[$BLOCKLIST_ARR]}"
|
||||
_bl_id_type="${_ID_TYPE[$BLOCKLIST_ARR]}"
|
||||
_bl_name_field="${_NAME[$BLOCKLIST_ARR]}"
|
||||
_bl_config_xml="${DOCKER_APPDATA_BASE}/${BLOCKLIST_ARR^}/config.xml"
|
||||
_bl_url="" _bl_key=""
|
||||
case "$BLOCKLIST_ARR" in
|
||||
lidarr) _bl_url="${LIDARR_URL:-}"; _bl_key="${LIDARR_API_KEY:-}" ;;
|
||||
@@ -419,8 +438,8 @@ if [[ -n "$BLOCKLIST_ACTION" ]]; then
|
||||
# Remove from all remote arrs
|
||||
for _bl_node_id in "${REMOTE_NODES[@]}"; do
|
||||
_bl_node_name="${!_bl_node_id}"
|
||||
_bl_result=$(_delete_remote_item "$_bl_node_id" "$_bl_port" "$_bl_ver" "$_bl_ep" \
|
||||
"$_bl_config_xml" "$_bl_id_field" "$_bl_id_type" "$BLOCKLIST_ID")
|
||||
_bl_result=$(_delete_remote_item "$_bl_node_id" "$_bl_port" "$_bl_ver" \
|
||||
"$BLOCKLIST_ARR" "$_bl_ep" "$_bl_id_field" "$_bl_id_type" "$BLOCKLIST_ID")
|
||||
case "$_bl_result" in
|
||||
200) log "Removed from $_bl_node_name ${BLOCKLIST_ARR^}: $_bl_display_name" ;;
|
||||
not_found) log "Not found on $_bl_node_name ${BLOCKLIST_ARR^} — already removed or not tracked" ;;
|
||||
@@ -459,7 +478,7 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo " Arr Port URL"
|
||||
for _arr in "${ARR_TYPES[@]}"; do
|
||||
local _url="" _key=""
|
||||
_url=""
|
||||
case "$_arr" in
|
||||
lidarr) _url="${LIDARR_URL:-not configured}" ;;
|
||||
sonarr) _url="${SONARR_URL:-not configured}" ;;
|
||||
@@ -478,19 +497,21 @@ fi
|
||||
|
||||
_resolve_node_ip() {
|
||||
local node_id="$1"
|
||||
local node_name="${!node_id}"
|
||||
local ts_name="${node_name,,}"
|
||||
local ip
|
||||
ip=$(tailscale ip -4 "$ts_name" 2>/dev/null)
|
||||
[[ -z "$ip" ]] && ip=$(tailscale status 2>/dev/null | \
|
||||
awk -v n="$ts_name" '$2 ~ "^" n { print $1; exit }')
|
||||
[[ -z "$ip" ]] && return 1
|
||||
echo "$ip"
|
||||
resolve_tailscale_ip "${!node_id}"
|
||||
}
|
||||
|
||||
# Check if arr is reachable on remote node
|
||||
# Args: node_id node_ip port api_ver arr_type
|
||||
_remote_arr_up() {
|
||||
local node_ip="$1" port="$2" api_ver="$3" config_xml="$4"
|
||||
local node_id="$1" node_ip="$2" port="$3" api_ver="$4" arr_type="$5"
|
||||
local _kvar="${node_id}_${arr_type^^}_API_KEY"; local cached_key="${!_kvar:-}"
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
curl -sf --max-time 5 \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
"http://${node_ip}:${port}/api/${api_ver}/system/status" >/dev/null 2>/dev/null
|
||||
return
|
||||
fi
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/${arr_type^}/config.xml"
|
||||
ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" \
|
||||
"KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
@@ -500,8 +521,17 @@ _remote_arr_up() {
|
||||
}
|
||||
|
||||
# Fetch full library from remote arr — returns raw JSON array
|
||||
# Args: node_id node_ip port api_ver arr_type endpoint
|
||||
_remote_library() {
|
||||
local node_ip="$1" port="$2" api_ver="$3" endpoint="$4" config_xml="$5"
|
||||
local node_id="$1" node_ip="$2" port="$3" api_ver="$4" arr_type="$5" endpoint="$6"
|
||||
local _kvar="${node_id}_${arr_type^^}_API_KEY"; local cached_key="${!_kvar:-}"
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
curl -sf --max-time "$ARR_SYNC_API_TIMEOUT" \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
"http://${node_ip}:${port}/api/${api_ver}/${endpoint}" 2>/dev/null
|
||||
return
|
||||
fi
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/${arr_type^}/config.xml"
|
||||
ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" \
|
||||
"KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
@@ -512,12 +542,31 @@ _remote_library() {
|
||||
}
|
||||
|
||||
# Fetch remote arr defaults: qualityProfileId, rootFolderPath, metadataProfileId (Lidarr)
|
||||
# Args: node_id node_ip port api_ver arr_type
|
||||
_remote_defaults() {
|
||||
local node_ip="$1" port="$2" api_ver="$3" arr_type="$4" config_xml="$5"
|
||||
local node_id="$1" node_ip="$2" port="$3" api_ver="$4" arr_type="$5"
|
||||
local _kvar="${node_id}_${arr_type^^}_API_KEY"; local cached_key="${!_kvar:-}"
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
local base_url="http://${node_ip}:${port}/api/${api_ver}"
|
||||
local qp rf mp
|
||||
qp=$(curl -sf -H "X-Api-Key: $cached_key" "${base_url}/qualityprofile" 2>/dev/null | jq '.[0].id // 1')
|
||||
rf=$(curl -sf -H "X-Api-Key: $cached_key" "${base_url}/rootfolder" 2>/dev/null | jq -r '.[0].path // ""')
|
||||
[[ -z "$qp" ]] && return 1
|
||||
if [[ "$arr_type" == "lidarr" ]]; then
|
||||
mp=$(curl -sf -H "X-Api-Key: $cached_key" "${base_url}/metadataprofile" 2>/dev/null | \
|
||||
jq 'map(select(.name == "Standard")) | .[0].id // .[0].id // 1')
|
||||
jq -n --argjson qp "$qp" --arg rf "$rf" --argjson mp "$mp" \
|
||||
'{qualityProfileId: $qp, rootFolderPath: $rf, metadataProfileId: $mp}'
|
||||
else
|
||||
jq -n --argjson qp "$qp" --arg rf "$rf" \
|
||||
'{qualityProfileId: $qp, rootFolderPath: $rf}'
|
||||
fi
|
||||
return
|
||||
fi
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/${arr_type^}/config.xml"
|
||||
local meta_field=""
|
||||
[[ "$arr_type" == "lidarr" ]] && \
|
||||
meta_field=', metadataProfileId: ($mp | map(select(.name == "Standard")) | .[0].id // .[0].id // 1)'
|
||||
|
||||
ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" bash <<REMOTE 2>/dev/null
|
||||
KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
@@ -530,11 +579,23 @@ jq -n --argjson qp "\$QP" --argjson rf "\$RF" --argjson mp "\$MP" \
|
||||
REMOTE
|
||||
}
|
||||
|
||||
# Add item to remote arr — payload is base64-encoded to avoid SSH quoting issues
|
||||
# Add item to remote arr — payload is base64-encoded to avoid quoting issues
|
||||
# Args: node_id node_ip port api_ver arr_type endpoint encoded_payload
|
||||
_remote_add() {
|
||||
local node_ip="$1" port="$2" api_ver="$3" endpoint="$4" config_xml="$5"
|
||||
local encoded="$6" # base64-encoded JSON body
|
||||
|
||||
local node_id="$1" node_ip="$2" port="$3" api_ver="$4" arr_type="$5" endpoint="$6"
|
||||
local encoded="$7"
|
||||
local _kvar="${node_id}_${arr_type^^}_API_KEY"; local cached_key="${!_kvar:-}"
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
local body
|
||||
body=$(printf '%s' "$encoded" | base64 -d)
|
||||
curl -sf -o /dev/null -w '%{http_code}' -X POST \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body" \
|
||||
"http://${node_ip}:${port}/api/${api_ver}/${endpoint}" 2>/dev/null
|
||||
return
|
||||
fi
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/${arr_type^}/config.xml"
|
||||
ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" bash <<REMOTE 2>/dev/null
|
||||
KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
@@ -548,15 +609,76 @@ curl -sf -o /dev/null -w '%{http_code}' -X POST \
|
||||
REMOTE
|
||||
}
|
||||
|
||||
# Trigger a full library rescan on a remote arr — used after merge-run to force the arr
|
||||
# to accept what is now on disk as ground truth rather than chasing stale file versions.
|
||||
# Args: arr_type node_id node_ip port api_ver
|
||||
_trigger_rescan() {
|
||||
local arr_type="$1" node_id="$2" node_ip="$3" port="$4" api_ver="$5"
|
||||
local command_name
|
||||
case "$arr_type" in
|
||||
sonarr) command_name="RefreshSeries" ;;
|
||||
radarr) command_name="RefreshMovie" ;;
|
||||
lidarr) command_name="RefreshArtist" ;;
|
||||
*) return 0 ;;
|
||||
esac
|
||||
|
||||
local encoded
|
||||
encoded=$(printf '{"name":"%s"}' "$command_name" | base64 -w0)
|
||||
local _kvar="${node_id}_${arr_type^^}_API_KEY"
|
||||
local cached_key="${!_kvar:-}"
|
||||
|
||||
local http_code
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
local body
|
||||
body=$(printf '%s' "$encoded" | base64 -d)
|
||||
http_code=$(curl -sf -o /dev/null -w '%{http_code}' -X POST \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body" \
|
||||
"http://${node_ip}:${port}/api/${api_ver}/command" 2>/dev/null)
|
||||
else
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/${arr_type^}/config.xml"
|
||||
http_code=$(ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" bash <<REMOTE 2>/dev/null
|
||||
KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
[[ -z "\$KEY" ]] && exit 1
|
||||
BODY=\$(printf '%s' '${encoded}' | base64 -d)
|
||||
curl -sf -o /dev/null -w '%{http_code}' -X POST \
|
||||
-H "X-Api-Key: \$KEY" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "\$BODY" \
|
||||
"http://localhost:${port}/api/${api_ver}/command"
|
||||
REMOTE
|
||||
)
|
||||
fi
|
||||
|
||||
if [[ "$http_code" == "201" || "$http_code" == "200" ]]; then
|
||||
log "${arr_type^}: triggered ${command_name} on ${node_id}"
|
||||
else
|
||||
warn "${arr_type^}: failed to trigger ${command_name} on ${node_id} (HTTP ${http_code:-timeout})"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── LOCAL HELPERS ─────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
_local_library() {
|
||||
local url="$1" api_key="$2" api_ver="$3" endpoint="$4"
|
||||
curl -sf --max-time "$ARR_SYNC_API_TIMEOUT" \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_ver}/${endpoint}" 2>/dev/null
|
||||
# endpoint doubles as arr_type here — artist/series/movie match lidarr/sonarr/radarr's
|
||||
# ARR_LIBRARY_ENDPOINT values exactly, so the same lookup works either direction.
|
||||
local arr_type
|
||||
case "$endpoint" in
|
||||
artist) arr_type="lidarr" ;;
|
||||
series) arr_type="sonarr" ;;
|
||||
movie) arr_type="radarr" ;;
|
||||
esac
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (kept
|
||||
# current every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back
|
||||
# to a live fetch when it's stale, and waits out an active rescan before either. This is
|
||||
# the LOCAL host's library only — the remote side of this sync isn't cached, since the
|
||||
# shared cache is per-host by design.
|
||||
arr_get_tracked_data "$arr_type" "$url" "$api_key" "$api_ver"
|
||||
}
|
||||
|
||||
_local_defaults() {
|
||||
@@ -643,6 +765,50 @@ _build_payload() {
|
||||
esac
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MONITORED ENFORCEMENT — re-monitor any unmonitored items via bulk editor ──────────────────
|
||||
# ==============================================================================================
|
||||
# All library members must be monitored. Unmonitored items won't be searched and eventually
|
||||
# lose their files when cleanup runs after the series is removed from the arr.
|
||||
|
||||
_enforce_monitored() {
|
||||
local arr_type="$1" url="$2" api_key="$3" api_ver="$4" local_json="$5"
|
||||
|
||||
local bulk_endpoint ids_key
|
||||
case "$arr_type" in
|
||||
sonarr) bulk_endpoint="series/editor"; ids_key="seriesIds" ;;
|
||||
radarr) bulk_endpoint="movie/editor"; ids_key="movieIds" ;;
|
||||
lidarr) bulk_endpoint="artist/editor"; ids_key="artistIds" ;;
|
||||
*) return 0 ;;
|
||||
esac
|
||||
|
||||
local ids_json count
|
||||
ids_json=$(echo "$local_json" | jq '[.[] | select(.monitored == false) | .id]' 2>/dev/null)
|
||||
count=$(echo "$ids_json" | jq 'length' 2>/dev/null || echo 0)
|
||||
[[ "$count" -eq 0 ]] && return 0
|
||||
|
||||
warn "${arr_type^}: $count unmonitored items — re-monitoring"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN: would re-monitor $count ${arr_type^} items"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local body http_code
|
||||
body=$(jq -n --arg k "$ids_key" --argjson ids "$ids_json" '{($k): $ids, monitored: true}')
|
||||
http_code=$(curl -sf -o /dev/null -w '%{http_code}' -X PUT \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body" \
|
||||
"${url}/api/${api_ver}/${bulk_endpoint}" 2>/dev/null)
|
||||
|
||||
if [[ "$http_code" == "200" || "$http_code" == "202" ]]; then
|
||||
echo "${arr_type^}: re-monitored $count items ✅"
|
||||
else
|
||||
warn "${arr_type^}: bulk re-monitor failed (HTTP $http_code)"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── CORE SYNC — one arr type across all remote nodes ──────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
@@ -655,7 +821,6 @@ _sync_arr() {
|
||||
local id_field="${_ID[$arr_type]}"
|
||||
local name_field="${_NAME[$arr_type]}"
|
||||
local id_type="${_ID_TYPE[$arr_type]}"
|
||||
local config_xml="${DOCKER_APPDATA_BASE}/$(echo "${arr_type^}")/config.xml"
|
||||
|
||||
# Resolve local credentials
|
||||
local local_url local_key
|
||||
@@ -696,6 +861,9 @@ _sync_arr() {
|
||||
|
||||
log "${arr_type^}: $local_count items in local library"
|
||||
|
||||
# ── Enforce monitored — fix any unmonitored items before sync ─────────────────────────────
|
||||
_enforce_monitored "$arr_type" "$local_url" "$local_key" "$api_ver" "$local_json"
|
||||
|
||||
local total_added_local=0 total_added_remote=0 total_skipped=0
|
||||
|
||||
# ── Sync with each remote node ─────────────────────────────────────────────────────────────
|
||||
@@ -709,13 +877,13 @@ _sync_arr() {
|
||||
continue
|
||||
}
|
||||
|
||||
if ! _remote_arr_up "$node_ip" "$port" "$api_ver" "$config_xml"; then
|
||||
if ! _remote_arr_up "$node_id" "$node_ip" "$port" "$api_ver" "$arr_type"; then
|
||||
log "${arr_type^}: not reachable on $node_name — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
local remote_json
|
||||
remote_json=$(_remote_library "$node_ip" "$port" "$api_ver" "$endpoint" "$config_xml")
|
||||
remote_json=$(_remote_library "$node_id" "$node_ip" "$port" "$api_ver" "$arr_type" "$endpoint")
|
||||
if [[ -z "$remote_json" ]] || ! echo "$remote_json" | jq -e '.' >/dev/null 2>&1; then
|
||||
warn "${arr_type^}: could not fetch library from $node_name — skipping"
|
||||
continue
|
||||
@@ -737,45 +905,50 @@ _sync_arr() {
|
||||
log "${arr_type^}: $remote_count items on $node_name"
|
||||
|
||||
# ── Remote → Local: items on remote not in local ───────────────────────────────────────
|
||||
local to_add_local=()
|
||||
for stable_id in "${!remote_ids[@]}"; do
|
||||
[[ -n "${local_ids[$stable_id]:-}" ]] && continue
|
||||
if _is_blocklisted "$arr_type" "$stable_id"; then
|
||||
log "BLOCKLISTED [$arr_type] $stable_id — skipping"
|
||||
(( total_skipped++ ))
|
||||
continue
|
||||
fi
|
||||
to_add_local+=("$stable_id")
|
||||
done
|
||||
# Skipped in force mode — local arr is authoritative; remote tracking does not propagate back.
|
||||
if [[ "$FORCE_MODE" == false ]]; then
|
||||
local to_add_local=()
|
||||
for stable_id in "${!remote_ids[@]}"; do
|
||||
[[ -n "${local_ids[$stable_id]:-}" ]] && continue
|
||||
if _is_blocklisted "$arr_type" "$stable_id"; then
|
||||
log "BLOCKLISTED [$arr_type] $stable_id — skipping"
|
||||
(( total_skipped++ ))
|
||||
continue
|
||||
fi
|
||||
to_add_local+=("$stable_id")
|
||||
done
|
||||
|
||||
if [[ "${#to_add_local[@]}" -gt 0 ]]; then
|
||||
local local_defs
|
||||
local_defs=$(_local_defaults "$local_url" "$local_key" "$api_ver" "$arr_type")
|
||||
if [[ -z "$local_defs" ]]; then
|
||||
warn "${arr_type^}: could not fetch local defaults — skipping adds from $node_name"
|
||||
else
|
||||
for stable_id in "${to_add_local[@]}"; do
|
||||
IFS='|' read -r display_name monitored <<< "${remote_ids[$stable_id]}"
|
||||
local payload
|
||||
payload=$(_build_payload "$arr_type" "$stable_id" "$display_name" \
|
||||
"$monitored" "$local_defs")
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
log "DRY RUN: would add to local ${arr_type^}: $display_name ($stable_id)"
|
||||
(( total_added_local++ ))
|
||||
else
|
||||
local http_code
|
||||
http_code=$(_local_add "$local_url" "$local_key" "$api_ver" \
|
||||
"$endpoint" "$payload")
|
||||
if [[ "$http_code" == "201" ]] || [[ "$http_code" == "200" ]]; then
|
||||
log "Added to local ${arr_type^}: $display_name"
|
||||
local_ids["$stable_id"]="${display_name}|${monitored}"
|
||||
if [[ "${#to_add_local[@]}" -gt 0 ]]; then
|
||||
local local_defs
|
||||
local_defs=$(_local_defaults "$local_url" "$local_key" "$api_ver" "$arr_type")
|
||||
if [[ -z "$local_defs" ]]; then
|
||||
warn "${arr_type^}: could not fetch local defaults — skipping adds from $node_name"
|
||||
else
|
||||
for stable_id in "${to_add_local[@]}"; do
|
||||
IFS='|' read -r display_name _ <<< "${remote_ids[$stable_id]}"
|
||||
local payload
|
||||
payload=$(_build_payload "$arr_type" "$stable_id" "$display_name" \
|
||||
"true" "$local_defs")
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
log "DRY RUN: would add to local ${arr_type^}: $display_name ($stable_id)"
|
||||
(( total_added_local++ ))
|
||||
else
|
||||
warn "Failed to add to local ${arr_type^}: $display_name (HTTP $http_code)"
|
||||
local http_code
|
||||
http_code=$(_local_add "$local_url" "$local_key" "$api_ver" \
|
||||
"$endpoint" "$payload")
|
||||
if [[ "$http_code" == "201" ]] || [[ "$http_code" == "200" ]]; then
|
||||
log "Added to local ${arr_type^}: $display_name"
|
||||
local_ids["$stable_id"]="${display_name}|true"
|
||||
(( total_added_local++ ))
|
||||
else
|
||||
warn "Failed to add to local ${arr_type^}: $display_name (HTTP $http_code)"
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
done
|
||||
done
|
||||
fi
|
||||
fi
|
||||
else
|
||||
log "${arr_type^}: force mode — skipping remote→local (local is authoritative)"
|
||||
fi
|
||||
|
||||
# ── Local → Remote: items on local not on remote ───────────────────────────────────────
|
||||
@@ -791,23 +964,23 @@ _sync_arr() {
|
||||
|
||||
if [[ "${#to_add_remote[@]}" -gt 0 ]]; then
|
||||
local remote_defs
|
||||
remote_defs=$(_remote_defaults "$node_ip" "$port" "$api_ver" "$arr_type" "$config_xml")
|
||||
remote_defs=$(_remote_defaults "$node_id" "$node_ip" "$port" "$api_ver" "$arr_type")
|
||||
if [[ -z "$remote_defs" ]]; then
|
||||
warn "${arr_type^}: could not fetch defaults from $node_name — skipping remote adds"
|
||||
else
|
||||
for stable_id in "${to_add_remote[@]}"; do
|
||||
IFS='|' read -r display_name monitored <<< "${local_ids[$stable_id]}"
|
||||
IFS='|' read -r display_name _ <<< "${local_ids[$stable_id]}"
|
||||
local payload
|
||||
payload=$(_build_payload "$arr_type" "$stable_id" "$display_name" \
|
||||
"$monitored" "$remote_defs")
|
||||
"true" "$remote_defs")
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
log "DRY RUN: would add to $node_name ${arr_type^}: $display_name ($stable_id)"
|
||||
(( total_added_remote++ ))
|
||||
else
|
||||
local encoded http_code
|
||||
encoded=$(printf '%s' "$payload" | base64 -w0)
|
||||
http_code=$(_remote_add "$node_ip" "$port" "$api_ver" "$endpoint" \
|
||||
"$config_xml" "$encoded")
|
||||
http_code=$(_remote_add "$node_id" "$node_ip" "$port" "$api_ver" \
|
||||
"$arr_type" "$endpoint" "$encoded")
|
||||
if [[ "$http_code" == "201" ]] || [[ "$http_code" == "200" ]]; then
|
||||
log "Added to $node_name ${arr_type^}: $display_name"
|
||||
(( total_added_remote++ ))
|
||||
@@ -819,9 +992,22 @@ _sync_arr() {
|
||||
fi
|
||||
fi
|
||||
|
||||
log " $node_name: +${#to_add_local[@]} local | +${#to_add_remote[@]} remote | $total_skipped blocklisted"
|
||||
# ── Force mode: trigger library rescan on remote ──────────────────────────────────────────
|
||||
# After files are settled by merge-run and tracking pushed by force arr-sync, the remote
|
||||
# arr rescans its library so it accepts the authoritative file versions as ground truth.
|
||||
if [[ "$FORCE_MODE" == true ]]; then
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
log "DRY RUN: would trigger ${arr_type^} rescan on $node_name"
|
||||
else
|
||||
_trigger_rescan "$arr_type" "$node_id" "$node_ip" "$port" "$api_ver"
|
||||
fi
|
||||
fi
|
||||
|
||||
unset remote_ids
|
||||
local _n_local=${#to_add_local[@]}
|
||||
local _n_remote=${#to_add_remote[@]}
|
||||
log " $node_name: +${_n_local} local | +${_n_remote} remote | $total_skipped blocklisted"
|
||||
|
||||
unset remote_ids to_add_local to_add_remote _n_local _n_remote
|
||||
declare -A remote_ids
|
||||
done
|
||||
|
||||
@@ -850,7 +1036,8 @@ echo "━━━━━ $ICON_SUMMARY ARR SYNC SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Node: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Peers: ${REMOTE_NODES[*]}"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes were made"
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes were made"
|
||||
[[ "$FORCE_MODE" == true ]] && echo " Mode: force (local authoritative — remote→local skipped, rescan triggered)"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
exit 0
|
||||
Executable
+794
@@ -0,0 +1,794 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================= Arrs Failed / Stalled Recovery =====================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Detect and recover failed imports and stalled downloads across Sonarr, Radarr,
|
||||
# and Lidarr. Blocklists the bad release and triggers a re-search — hands-free
|
||||
# overnight recovery.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Five problem types detected from the arr queue API:
|
||||
# importFailed — downloaded but arr couldn't import the file
|
||||
# importPending — downloaded, stuck waiting to import (will not self-resolve)
|
||||
# importBlocked — downloaded, but arr matched the release to the wrong media
|
||||
# by grab-history ID instead of by title and refuses to import
|
||||
# (permanent block, never self-resolves)
|
||||
# error status — serious failure not covered by the above two states
|
||||
# stalled — download stuck with no connections or no progress
|
||||
#
|
||||
# Never touches items with state "downloading" or "imported" — safe to run anytime.
|
||||
# Items newer than ARR_IMPORT_RECOVERY_AGE are skipped — gives arr time to retry first.
|
||||
#
|
||||
# importBlocked items get one extra check first (try_smart_import, Sonarr/Radarr only):
|
||||
# most are junk/duplicates and fall straight through to the normal 3-step response below,
|
||||
# but some are a release arr already correctly parsed — episode/movie identified, quality
|
||||
# and language known — that's just tripping the title-vs-grab-history safety net. If the
|
||||
# target has no file yet (missing) or the candidate is a same-language resolution upgrade
|
||||
# over what's already there, it's imported directly instead of being discarded. See
|
||||
# ARR_SMART_IMPORT_ENABLED in CONFIGURATION below.
|
||||
#
|
||||
# Per problem item that isn't smart-imported (3-step response):
|
||||
# 1. Blocklist the release — prevents re-grabbing the same bad release
|
||||
# 2. Remove from queue — cleans up the failed item
|
||||
# 3. Trigger new search — finds a different release automatically
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Hands-Free Recovery
|
||||
# The script completes the full recovery cycle autonomously — blocklist, remove,
|
||||
# re-search. No operator decision required. A failed import at midnight resolves
|
||||
# itself before morning without any intervention.
|
||||
#
|
||||
# Smart Import Is Conservative By Design
|
||||
# try_smart_import only acts when every file in the download is unambiguous: no
|
||||
# rejections from arr's own analysis, and (no existing file) or (matching language
|
||||
# plus a strictly higher resolution). Any ambiguity — mixed multi-episode files,
|
||||
# unknown language, equal-or-lower quality, wrong language — falls straight through
|
||||
# to blocklist+research, exactly today's behavior. It only ever adds a chance to
|
||||
# keep something worth keeping; it never makes the no-smart-import case worse.
|
||||
#
|
||||
# Age Gate Before Action
|
||||
# Items newer than ARR_IMPORT_RECOVERY_AGE are skipped. Arrs have their own
|
||||
# retry logic — acting immediately would race against it. The age gate gives
|
||||
# the arr time to self-resolve before this script escalates.
|
||||
#
|
||||
# Blocklist First
|
||||
# The bad release is blocklisted before removal and re-search. Without this,
|
||||
# the re-search can re-grab the same release that just failed.
|
||||
#
|
||||
# Circuit Breaker Per Media Item
|
||||
# Some items can never resolve via blind retry — e.g. an album missing 1-2
|
||||
# tracks where every available release is a different edition that doesn't
|
||||
# match. Without a limit, the same media ID gets blocklisted + re-searched
|
||||
# forever, every run, burning bandwidth and indexer queries for nothing.
|
||||
# After ARR_RECOVERY_MAX_ATTEMPTS consecutive failures for the same
|
||||
# (arr_type, media_id), the item is still blocklisted/cleaned from the queue
|
||||
# but search is no longer auto-triggered — it's flagged chronic and left for
|
||||
# manual review instead.
|
||||
#
|
||||
# "Consecutive" is enforced, not just counted — a media_id's failure count
|
||||
# is pruned at the end of every process_arr() pass if it no longer appears
|
||||
# in that run's problem-item set (2026-07-19 fix: counts were never reset on
|
||||
# success, so an item that failed a few times months apart and then imported
|
||||
# fine could still get stuck permanently chronic from stale history).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# acquire_lock — prevents concurrent runs overlapping
|
||||
# jq validation — exits if jq not installed (required for JSON parsing)
|
||||
# API pre-flight — checks each arr is reachable before querying queue
|
||||
# Version check — check_arr_version() verifies running arr matches master.conf major
|
||||
# version; exits rather than silently misoperating after upgrade
|
||||
# Age threshold — skips items newer than ARR_IMPORT_RECOVERY_AGE (default 6hr)
|
||||
# Silent by default — only problems produce output, clean arrs stay silent
|
||||
#
|
||||
# API version mapping (endpoint paths differ from major version labels):
|
||||
# Sonarr v4 → /api/v3/ (v3 endpoint retained in v4)
|
||||
# Radarr v6 → /api/v3/ (v3 endpoint retained in v6)
|
||||
# Lidarr v3 → /api/v1/ (different from Sonarr/Radarr)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ARR_RECOVERY_STATS — stats file written after each run (read by coffee report)
|
||||
# ARR_RECOVERY_FAILURE_COUNTS — per (arr_type, media_id) consecutive-failure counts,
|
||||
# persists across runs so the circuit breaker survives restarts.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_SONARR_URL / HOST*_SONARR_API_KEY / HOST*_SONARR_RECOVERY
|
||||
# HOST*_RADARR_URL / HOST*_RADARR_API_KEY / HOST*_RADARR_RECOVERY
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY / HOST1_LIDARR_RECOVERY
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# ARR_IMPORT_RECOVERY_AGE — hours before item is eligible for recovery (default: 6)
|
||||
# ARR_RECOVERY_MAX_ATTEMPTS — consecutive failures before an item is flagged chronic
|
||||
# and auto re-search stops (default: 3)
|
||||
# ARR_SMART_IMPORT_ENABLED — try_smart_import gate for importBlocked items, Sonarr/
|
||||
# Radarr only — Lidarr's manual-import matching doesn't
|
||||
# reliably resolve album/track context (default: true)
|
||||
# ARR_SMART_IMPORT_PREFERRED_LANGUAGE — only import as a match/upgrade if the
|
||||
# candidate is this language; existing files in a
|
||||
# different language are always treated as upgradeable
|
||||
# (default: English)
|
||||
# SONARR_VERSION_MAJOR — expected Sonarr major version (e.g. 4)
|
||||
# RADARR_VERSION_MAJOR — expected Radarr major version (e.g. 6)
|
||||
# LIDARR_VERSION_MAJOR — expected Lidarr major version (e.g. 3)
|
||||
# ARR_RECOVERY_STATS — stats file path (read by coffee report)
|
||||
# ARR_RECOVERY_FAILURE_COUNTS — failure-count state file path
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arrs_failed_stalled_recovery.sh — normal run
|
||||
# arrs_failed_stalled_recovery.sh --dry-run — show what would be actioned, no changes
|
||||
# arrs_failed_stalled_recovery.sh --log — verbose output
|
||||
# arrs_failed_stalled_recovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: 0 5 * * * (5am daily)
|
||||
# Or every 6hr: 0 */6 * * * (matches ARR_IMPORT_RECOVERY_AGE default)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases SONARR_*, RADARR_*, LIDARR_* vars
|
||||
detect_hosts
|
||||
|
||||
# jq is required — not optional — for JSON parsing
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq is not installed — required for arr API JSON parsing"
|
||||
error "Install: apt-get install jq or brew install jq"
|
||||
notify "arrs_failed_stalled_recovery failed on $(hostname) — jq not installed" \
|
||||
"Arr Recovery" "warning"
|
||||
exit 1
|
||||
fi
|
||||
log "jq found"
|
||||
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no items will be blocklisted or searched"
|
||||
|
||||
# Age threshold in seconds
|
||||
AGE_THRESHOLD_SECONDS=$(( ARR_IMPORT_RECOVERY_AGE * 3600 ))
|
||||
ARR_RECOVERY_MAX_ATTEMPTS="${ARR_RECOVERY_MAX_ATTEMPTS:-3}"
|
||||
ARR_RECOVERY_FAILURE_COUNTS="${ARR_RECOVERY_FAILURE_COUNTS:-$DATA_DIR/arr_recovery_failure_counts.db}"
|
||||
|
||||
log "$ICON_GEAR Config: age-threshold=${ARR_IMPORT_RECOVERY_AGE}hr max-attempts=${ARR_RECOVERY_MAX_ATTEMPTS} sonarr-v${SONARR_VERSION_MAJOR} radarr-v${RADARR_VERSION_MAJOR} lidarr-v${LIDARR_VERSION_MAJOR:-?}"
|
||||
|
||||
# Load persisted per-item failure counts — key is "arr_type:media_id"
|
||||
declare -A FAILURE_COUNTS
|
||||
if [[ -f "$ARR_RECOVERY_FAILURE_COUNTS" ]]; then
|
||||
while IFS='|' read -r _key _count; do
|
||||
[[ -z "$_key" ]] && continue
|
||||
FAILURE_COUNTS["$_key"]="$_count"
|
||||
done < "$ARR_RECOVERY_FAILURE_COUNTS"
|
||||
fi
|
||||
|
||||
TOTAL_ACTIONED=0
|
||||
TOTAL_SKIPPED=0
|
||||
TOTAL_CHRONIC=0
|
||||
TOTAL_SMART_IMPORTED=0
|
||||
ARR_SUMMARIES=()
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Sonarr: ${SONARR_URL:-not configured} (recovery: ${SONARR_RECOVERY:-true})"
|
||||
echo "$ICON_SYNC Radarr: ${RADARR_URL:-not configured} (recovery: ${RADARR_RECOVERY:-true})"
|
||||
echo "$ICON_SYNC Lidarr: ${LIDARR_URL:-not configured on this host} (recovery: ${LIDARR_RECOVERY:-false})"
|
||||
echo "$ICON_TIME Age thresh: ${ARR_IMPORT_RECOVERY_AGE}hr"
|
||||
echo "$ICON_GEAR Max attempts: ${ARR_RECOVERY_MAX_ATTEMPTS:-3} (chronic after this many)"
|
||||
echo "$ICON_GEAR Smart import: ${ARR_SMART_IMPORT_ENABLED:-true} (preferred language: ${ARR_SMART_IMPORT_PREFERRED_LANGUAGE:-English})"
|
||||
echo "$ICON_GEAR Sonarr ver: v${SONARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Radarr ver: v${RADARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Lidarr ver: v${LIDARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_NOTIFY Notify: unRAID=${NOTIFY_UNRAID:-false} Discord=$([[ -n "${MY_DISCORD_WEBHOOK:-}" ]] && echo enabled || echo disabled)"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Check if a queue item is older than ARR_IMPORT_RECOVERY_AGE
|
||||
# Returns 0 (old enough) or 1 (too new — skip)
|
||||
item_is_old_enough() {
|
||||
local added="$1"
|
||||
[[ -z "$added" ]] && return 0 # no date = treat as old enough, safe to act
|
||||
local added_epoch
|
||||
added_epoch=$(date -d "$added" +%s 2>/dev/null) || return 0
|
||||
local age_seconds=$(( $(date +%s) - added_epoch ))
|
||||
[[ "$age_seconds" -ge "$AGE_THRESHOLD_SECONDS" ]]
|
||||
}
|
||||
|
||||
# Query the arr queue API and return all records, paginated.
|
||||
# A single page=1&pageSize=200 request silently misses everything past record 200 —
|
||||
# on a busy Sonarr instance the queue can run into the thousands (e.g. a large
|
||||
# missing-episode search campaign), which pushed every importBlocked/warning item
|
||||
# past page 1 and made this whole script blind to them despite matching correctly.
|
||||
# Args: url, api_key, api_version
|
||||
get_queue_data() {
|
||||
local url="$1" api_key="$2" api_version="$3"
|
||||
local page=1 page_size=250 max_pages=50
|
||||
local page_data page_count
|
||||
# Accumulate pages as files rather than growing a shell variable — on a large
|
||||
# queue (thousands of records) passing the combined JSON through --argjson
|
||||
# blows past ARG_MAX ("Argument list too long"). jq -s reads files instead.
|
||||
local tmp_dir
|
||||
tmp_dir=$(mktemp -d)
|
||||
trap 'rm -rf "$tmp_dir"' RETURN
|
||||
|
||||
while [[ "$page" -le "$max_pages" ]]; do
|
||||
page_data=$(curl -sf --max-time 15 \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/queue?page=${page}&pageSize=${page_size}&includeUnknownSeriesItems=true&includeUnknownArtistItems=true" \
|
||||
2>/dev/null)
|
||||
[[ -z "$page_data" ]] && break
|
||||
|
||||
page_count=$(echo "$page_data" | jq '.records // [] | length' 2>/dev/null)
|
||||
[[ -z "$page_count" || "$page_count" -eq 0 ]] && break
|
||||
|
||||
echo "$page_data" | jq -c '.records // []' > "$tmp_dir/page_${page}.json"
|
||||
|
||||
[[ "$page_count" -lt "$page_size" ]] && break
|
||||
(( page++ ))
|
||||
done
|
||||
|
||||
jq -c -s '{totalRecords: ([.[][]] | length), records: [.[][]]}' "$tmp_dir"/page_*.json 2>/dev/null \
|
||||
|| echo '{"totalRecords":0,"records":[]}'
|
||||
}
|
||||
|
||||
# Blocklist and remove a queue item
|
||||
# Args: url, api_key, api_version, queue_id
|
||||
blocklist_item() {
|
||||
local url="$1" api_key="$2" api_version="$3" queue_id="$4"
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would blocklist queue item $queue_id"
|
||||
return 0
|
||||
fi
|
||||
curl -sf --max-time 15 \
|
||||
-X DELETE \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/queue/${queue_id}?removeFromClient=true&blocklist=true&skipRedownload=false" \
|
||||
>/dev/null 2>&1
|
||||
}
|
||||
|
||||
# Trigger a new search for the media item
|
||||
# Args: url, api_key, api_version, arr_type, media_id
|
||||
trigger_search() {
|
||||
local url="$1" api_key="$2" api_version="$3" arr_type="$4" media_id="$5"
|
||||
local command body
|
||||
case "$arr_type" in
|
||||
sonarr) command="EpisodeSearch"; body="{\"name\":\"EpisodeSearch\",\"episodeIds\":[$media_id]}" ;;
|
||||
radarr) command="MoviesSearch"; body="{\"name\":\"MoviesSearch\",\"movieIds\":[$media_id]}" ;;
|
||||
lidarr) command="AlbumSearch"; body="{\"name\":\"AlbumSearch\",\"albumIds\":[$media_id]}" ;;
|
||||
esac
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would trigger $command for media ID $media_id"
|
||||
return 0
|
||||
fi
|
||||
curl -sf --max-time 15 \
|
||||
-X POST \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body" \
|
||||
"${url}/api/${api_version}/command" \
|
||||
>/dev/null 2>&1
|
||||
}
|
||||
|
||||
# Decide whether an importBlocked download is actually worth keeping, and import
|
||||
# it directly if so — instead of always discarding it via blocklist+research.
|
||||
#
|
||||
# Fires the ManualImport command and returns as soon as it's accepted (HTTP 201)
|
||||
# rather than polling for completion. Deliberately does NOT fall back to
|
||||
# blocklist_item() after a successful trigger: Radarr/Sonarr's import runs async
|
||||
# in the background, and blocklisting (which deletes the source via
|
||||
# removeFromClient) right after firing it would race a still-in-progress import
|
||||
# for anything but the smallest files. If the import silently fails, the item
|
||||
# simply reappears as importBlocked next run and gets tried again — safe, if not
|
||||
# maximally fast, since the source file is never touched by this function.
|
||||
#
|
||||
# Returns 0 if a smart-import was triggered (caller should skip the normal
|
||||
# blocklist+research path for this item), 1 if declined or failed (caller should
|
||||
# fall through to the normal path exactly as before this function existed).
|
||||
# Args: url, api_key, api_version, arr_type, download_id, title
|
||||
try_smart_import() {
|
||||
local url="$1" api_key="$2" api_version="$3" arr_type="$4" download_id="$5" title="$6"
|
||||
|
||||
# Lidarr's manual-import matching doesn't reliably resolve album/track context
|
||||
# (confirmed 2026-07-15 — 659/659 track candidates came back with no album
|
||||
# match at all) — not worth attempting, always fall through to normal handling.
|
||||
[[ "$arr_type" == "lidarr" ]] && return 1
|
||||
[[ -z "$download_id" ]] && return 1
|
||||
|
||||
local candidates
|
||||
candidates=$(curl -sf --max-time 30 \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/manualimport?downloadId=${download_id}" \
|
||||
2>/dev/null)
|
||||
[[ -z "$candidates" || "$candidates" == "[]" || "$candidates" == "null" ]] && return 1
|
||||
|
||||
# Any rejected file (Sample, Unknown Movie/Series, "Not an upgrade", etc.)
|
||||
# disqualifies the whole download — conservative by design.
|
||||
local rejected_count
|
||||
rejected_count=$(echo "$candidates" | jq '[.[] | select(.rejections | length > 0)] | length' 2>/dev/null)
|
||||
[[ -z "$rejected_count" || "$rejected_count" -gt 0 ]] && return 1
|
||||
|
||||
local file_count
|
||||
file_count=$(echo "$candidates" | jq 'length' 2>/dev/null)
|
||||
[[ -z "$file_count" || "$file_count" -eq 0 ]] && return 1
|
||||
|
||||
local preferred_lang="${ARR_SMART_IMPORT_PREFERRED_LANGUAGE:-English}"
|
||||
local qualifying_files=()
|
||||
local i entry target_id has_file existing existing_res existing_lang candidate_res candidate_lang decision
|
||||
|
||||
for (( i=0; i<file_count; i++ )); do
|
||||
entry=$(echo "$candidates" | jq -c ".[$i]" 2>/dev/null)
|
||||
[[ -z "$entry" ]] && return 1
|
||||
|
||||
candidate_res=$(echo "$entry" | jq -r '.quality.quality.resolution // 0' 2>/dev/null)
|
||||
candidate_lang=$(echo "$entry" | jq -r '.languages[0].name // "Unknown"' 2>/dev/null)
|
||||
existing_res=0
|
||||
existing_lang="Unknown"
|
||||
target_id=""
|
||||
has_file="false"
|
||||
|
||||
case "$arr_type" in
|
||||
sonarr)
|
||||
# Multi-episode files complicate the existing-quality comparison per
|
||||
# episode — skip rather than guess when a release covers more than one.
|
||||
[[ "$(echo "$entry" | jq '.episodes | length' 2>/dev/null)" != "1" ]] && return 1
|
||||
target_id=$(echo "$entry" | jq -r '.episodes[0].id // empty' 2>/dev/null)
|
||||
has_file=$(echo "$entry" | jq -r '.episodes[0].hasFile // false' 2>/dev/null)
|
||||
if [[ "$has_file" == "true" ]]; then
|
||||
existing=$(curl -sf --max-time 15 -H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/episode/${target_id}?includeEpisodeFile=true" 2>/dev/null)
|
||||
existing_res=$(echo "$existing" | jq -r '.episodeFile.quality.quality.resolution // 0' 2>/dev/null)
|
||||
existing_lang=$(echo "$existing" | jq -r '.episodeFile.languages[0].name // "Unknown"' 2>/dev/null)
|
||||
fi
|
||||
# /manualimport only nests the IDs under .series.id / .episodes[].id —
|
||||
# the ManualImport command body needs them flattened to top-level
|
||||
# seriesId/episodeIds or Sonarr rejects the whole command with
|
||||
# "Series with ID 0 does not exist" (confirmed live 2026-07-19: every
|
||||
# smart-import this run reported as successful had actually failed
|
||||
# this way, silently, since the caller only checks the HTTP 201 accept).
|
||||
entry=$(echo "$entry" | jq -c --argjson eid "$target_id" \
|
||||
'. + {seriesId: .series.id, episodeIds: [$eid]}' 2>/dev/null)
|
||||
;;
|
||||
radarr)
|
||||
target_id=$(echo "$entry" | jq -r '.movie.id // empty' 2>/dev/null)
|
||||
has_file=$(echo "$entry" | jq -r '.movie.hasFile // false' 2>/dev/null)
|
||||
if [[ "$has_file" == "true" ]]; then
|
||||
existing=$(echo "$entry" | jq -c '.movie.movieFile // empty' 2>/dev/null)
|
||||
if [[ -z "$existing" || "$existing" == "null" ]]; then
|
||||
existing=$(curl -sf --max-time 15 -H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/movie/${target_id}" 2>/dev/null | jq -c '.movieFile // empty')
|
||||
fi
|
||||
existing_res=$(echo "$existing" | jq -r '.quality.quality.resolution // 0' 2>/dev/null)
|
||||
existing_lang=$(echo "$existing" | jq -r '.languages[0].name // "Unknown"' 2>/dev/null)
|
||||
fi
|
||||
# Same flattening issue as Sonarr above — command body needs a
|
||||
# top-level movieId or Radarr rejects it with "Movie with ID 0
|
||||
# does not exist".
|
||||
entry=$(echo "$entry" | jq -c --argjson mid "$target_id" '. + {movieId: $mid}' 2>/dev/null)
|
||||
;;
|
||||
esac
|
||||
|
||||
[[ -z "$target_id" || -z "$entry" ]] && return 1
|
||||
|
||||
if [[ "$has_file" != "true" ]]; then
|
||||
decision="import" # nothing there yet — fills a real gap
|
||||
elif [[ "$candidate_lang" != "$preferred_lang" ]]; then
|
||||
decision="decline" # never replace anything with a non-preferred language
|
||||
elif [[ "$existing_lang" != "$preferred_lang" ]]; then
|
||||
decision="import" # existing is wrong-language, candidate is right — upgrade
|
||||
elif [[ "$candidate_res" -gt "$existing_res" ]]; then
|
||||
decision="import" # same language, strictly higher resolution — upgrade
|
||||
else
|
||||
decision="decline" # same or worse, same language — no benefit
|
||||
fi
|
||||
|
||||
[[ "$decision" == "decline" ]] && return 1
|
||||
qualifying_files+=("$entry")
|
||||
done
|
||||
|
||||
[[ "${#qualifying_files[@]}" -eq 0 ]] && return 1
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn " DRY RUN — would smart-import: $title"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local files_json cmd_body response http_code cmd_id cmd_status
|
||||
files_json=$(printf '%s\n' "${qualifying_files[@]}" | jq -s -c '.' 2>/dev/null)
|
||||
[[ -z "$files_json" ]] && return 1
|
||||
cmd_body=$(jq -c -n --argjson files "$files_json" \
|
||||
'{name:"ManualImport", files:$files, importMode:"auto"}' 2>/dev/null)
|
||||
[[ -z "$cmd_body" ]] && return 1
|
||||
|
||||
response=$(curl -s -w '\n%{http_code}' -X POST \
|
||||
-H "X-Api-Key: $api_key" -H "Content-Type: application/json" \
|
||||
-d "$cmd_body" \
|
||||
"${url}/api/${api_version}/command" 2>/dev/null)
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
cmd_id=$(echo "$response" | head -n -1 | jq -r '.id // empty' 2>/dev/null)
|
||||
|
||||
[[ "$http_code" != "201" || -z "$cmd_id" ]] && return 1
|
||||
|
||||
# A bad payload (e.g. the seriesId/movieId=0 bug this was written to catch)
|
||||
# fails in ~10ms — well before any real file copy would even start — so a
|
||||
# brief poll here catches that failure class without racing a genuinely
|
||||
# long-running import, which is the reason this doesn't poll to completion.
|
||||
local _i
|
||||
for _i in 1 2 3; do
|
||||
sleep 1
|
||||
cmd_status=$(curl -sf --max-time 10 -H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/command/${cmd_id}" 2>/dev/null | jq -r '.status // empty')
|
||||
[[ "$cmd_status" == "failed" ]] && return 1
|
||||
[[ "$cmd_status" == "completed" ]] && break
|
||||
done
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── PROCESS AN ARR ────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Args: display_name, arr_type, url, api_key, api_version, enabled,
|
||||
# version_major, version_api_prefix
|
||||
#
|
||||
# Exits cleanly if disabled.
|
||||
# Checks API reachability and version before touching queue.
|
||||
# Processes each problem item: blocklist + trigger new search.
|
||||
# Silent when clean — only warns when problems found or actioned.
|
||||
|
||||
process_arr() {
|
||||
local arr_name="$1"
|
||||
local arr_type="$2"
|
||||
local url="$3"
|
||||
local api_key="$4"
|
||||
local api_version="$5"
|
||||
local enabled="$6"
|
||||
local version_major="$7"
|
||||
local version_api_prefix="$8"
|
||||
|
||||
local actioned=0 skipped_new=0 chronic=0 smart_imported=0
|
||||
local is_chronic fail_key fail_count
|
||||
local -A seen_media_ids # media_ids appearing as a problem this run — anything NOT in
|
||||
# here by the end has stopped being a problem and has its
|
||||
# FAILURE_COUNTS entry pruned below
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC $arr_name ━━━"
|
||||
|
||||
# Disabled — skip cleanly
|
||||
if [[ "$enabled" != "true" ]]; then
|
||||
log "$arr_name recovery disabled — skipping"
|
||||
ARR_SUMMARIES+=("$arr_name: disabled")
|
||||
return
|
||||
fi
|
||||
|
||||
# URL not configured on this host — skip cleanly
|
||||
if [[ -z "$url" ]]; then
|
||||
log "$arr_name not configured on $MY_ID — skipping"
|
||||
ARR_SUMMARIES+=("$arr_name: not configured on $MY_ID")
|
||||
return
|
||||
fi
|
||||
|
||||
# API reachability
|
||||
if ! check_api "$url" "$arr_name" 10; then
|
||||
warn "$arr_name API unreachable — skipping"
|
||||
ARR_SUMMARIES+=("$arr_name: unreachable")
|
||||
return
|
||||
fi
|
||||
|
||||
# Version check — exit if API structure may have changed
|
||||
if ! check_arr_version "$url" "$api_key" "$version_api_prefix" \
|
||||
"$version_major" "$arr_name"; then
|
||||
ARR_SUMMARIES+=("$arr_name: version mismatch — skipped")
|
||||
return
|
||||
fi
|
||||
|
||||
# Fetch queue
|
||||
local queue_data
|
||||
queue_data=$(get_queue_data "$url" "$api_key" "$api_version")
|
||||
if [[ -z "$queue_data" ]]; then
|
||||
warn "$arr_name — could not retrieve queue data"
|
||||
ARR_SUMMARIES+=("$arr_name: queue fetch failed")
|
||||
return
|
||||
fi
|
||||
|
||||
local total_records
|
||||
total_records=$(echo "$queue_data" | jq '.totalRecords // 0' 2>/dev/null)
|
||||
log "$arr_name queue: $total_records total items"
|
||||
|
||||
# Filter for problem items — never touch downloading or imported
|
||||
local problem_items
|
||||
problem_items=$(echo "$queue_data" | jq -c '
|
||||
.records // [] |
|
||||
.[] |
|
||||
select(
|
||||
.trackedDownloadState != "downloading" and
|
||||
.trackedDownloadState != "imported" and
|
||||
(
|
||||
.trackedDownloadState == "importFailed" or
|
||||
.trackedDownloadState == "importPending" or
|
||||
.trackedDownloadState == "importBlocked" or
|
||||
.trackedDownloadStatus == "error" or
|
||||
(.status == "warning" and (
|
||||
(.errorMessage // "" | ascii_downcase | contains("stalled")) or
|
||||
(.statusMessages // [] | .[] | .messages // [] | .[] |
|
||||
ascii_downcase | contains("stalled"))
|
||||
))
|
||||
)
|
||||
)
|
||||
' 2>/dev/null)
|
||||
|
||||
if [[ -z "$problem_items" ]]; then
|
||||
echo "$arr_name — clean ✅ no failed imports or stalled downloads"
|
||||
ARR_SUMMARIES+=("$arr_name: clean ✅")
|
||||
return
|
||||
fi
|
||||
|
||||
local problem_count
|
||||
problem_count=$(echo "$problem_items" | wc -l)
|
||||
warn "$arr_name — found $problem_count problem item(s)"
|
||||
|
||||
# Process each problem item
|
||||
while IFS= read -r item; do
|
||||
[[ -z "$item" ]] && continue
|
||||
|
||||
local queue_id title added tracked_state tracked_status problem_type media_id download_id
|
||||
|
||||
queue_id=$(echo "$item" | jq -r '.id // empty' 2>/dev/null)
|
||||
title=$(echo "$item" | jq -r '.title // "Unknown"' 2>/dev/null)
|
||||
added=$(echo "$item" | jq -r '.added // empty' 2>/dev/null)
|
||||
tracked_state=$(echo "$item" | jq -r '.trackedDownloadState // ""' 2>/dev/null)
|
||||
tracked_status=$(echo "$item" | jq -r '.trackedDownloadStatus // ""' 2>/dev/null)
|
||||
|
||||
# Human-readable problem type
|
||||
case "$tracked_state" in
|
||||
importFailed) problem_type="import failed" ;;
|
||||
importPending) problem_type="import pending/stuck" ;;
|
||||
importBlocked) problem_type="import blocked (matched by ID)" ;;
|
||||
*)
|
||||
[[ "$tracked_status" == "error" ]] && \
|
||||
problem_type="error" || problem_type="stalled"
|
||||
;;
|
||||
esac
|
||||
|
||||
# Media ID for search trigger
|
||||
case "$arr_type" in
|
||||
sonarr) media_id=$(echo "$item" | jq -r '.episodeId // .episode.id // empty' 2>/dev/null) ;;
|
||||
radarr) media_id=$(echo "$item" | jq -r '.movieId // .movie.id // empty' 2>/dev/null) ;;
|
||||
lidarr) media_id=$(echo "$item" | jq -r '.albumId // .album.id // empty' 2>/dev/null) ;;
|
||||
esac
|
||||
|
||||
[[ -n "$media_id" ]] && seen_media_ids["$media_id"]=1
|
||||
|
||||
[[ -z "$queue_id" ]] && continue
|
||||
|
||||
# Age check — skip items that are too new to have self-resolved
|
||||
if ! item_is_old_enough "$added"; then
|
||||
log " Skipping (too new < ${ARR_IMPORT_RECOVERY_AGE}hr): $title"
|
||||
(( skipped_new++ ))
|
||||
(( TOTAL_SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# importBlocked gets one extra chance before the normal blocklist path —
|
||||
# most are junk/duplicates and fall straight through unchanged, but a
|
||||
# clean same-language upgrade or gap-fill gets imported directly instead
|
||||
# of discarded. See try_smart_import() for the full decision logic.
|
||||
if [[ "${ARR_SMART_IMPORT_ENABLED:-true}" == "true" ]] && \
|
||||
[[ "$tracked_state" == "importBlocked" ]]; then
|
||||
download_id=$(echo "$item" | jq -r '.downloadId // empty' 2>/dev/null)
|
||||
if try_smart_import "$url" "$api_key" "$api_version" "$arr_type" "$download_id" "$title"; then
|
||||
log " $ICON_DONE Smart-imported (upgrade/gap-fill): $title"
|
||||
(( smart_imported++ ))
|
||||
(( TOTAL_SMART_IMPORTED++ ))
|
||||
(( actioned++ ))
|
||||
(( TOTAL_ACTIONED++ ))
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
|
||||
warn " $ICON_TRASH $problem_type — $title"
|
||||
|
||||
# Step 1: Blocklist + remove from queue
|
||||
if ! blocklist_item "$url" "$api_key" "$api_version" "$queue_id"; then
|
||||
warn " Failed to blocklist: $title"
|
||||
(( TOTAL_SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
log " Blocklisted: $queue_id"
|
||||
|
||||
# Step 2: Circuit breaker — track consecutive failures per (arr_type, media_id).
|
||||
# Some items can never resolve via blind retry (e.g. an album missing 1-2 tracks
|
||||
# where no available release matches the existing edition) — without this, the
|
||||
# same item gets blocklisted + re-searched forever, every run.
|
||||
is_chronic=false
|
||||
if [[ -n "$media_id" ]]; then
|
||||
fail_key="${arr_type}:${media_id}"
|
||||
fail_count=$(( ${FAILURE_COUNTS[$fail_key]:-0} + 1 ))
|
||||
FAILURE_COUNTS[$fail_key]="$fail_count"
|
||||
if [[ "$fail_count" -gt "$ARR_RECOVERY_MAX_ATTEMPTS" ]]; then
|
||||
is_chronic=true
|
||||
(( chronic++ ))
|
||||
(( TOTAL_CHRONIC++ ))
|
||||
warn " Chronic (${fail_count} consecutive failures) — needs manual review: $title"
|
||||
[[ "$fail_count" -eq $(( ARR_RECOVERY_MAX_ATTEMPTS + 1 )) ]] && \
|
||||
notify "$arr_name item now chronic after ${ARR_RECOVERY_MAX_ATTEMPTS} failed attempts — needs manual review: $title" \
|
||||
"Arr Recovery" "warning"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Step 3: Trigger new search — skipped for chronic items
|
||||
if [[ "$is_chronic" == true ]]; then
|
||||
log " Skipping auto re-search (chronic): $title"
|
||||
elif [[ -n "$media_id" ]]; then
|
||||
if trigger_search "$url" "$api_key" "$api_version" "$arr_type" "$media_id"; then
|
||||
log " New search triggered: $title"
|
||||
else
|
||||
warn " Blocklisted but search trigger failed: $title"
|
||||
fi
|
||||
else
|
||||
warn " Blocklisted but no media ID found — search not triggered: $title"
|
||||
fi
|
||||
|
||||
(( actioned++ ))
|
||||
(( TOTAL_ACTIONED++ ))
|
||||
|
||||
done <<< "$problem_items"
|
||||
|
||||
# Prune stale failure counts — anything for this arr_type that isn't a problem in this
|
||||
# run's queue snapshot has either imported successfully or is otherwise no longer stuck.
|
||||
# FAILURE_COUNTS never decremented on success (confirmed live 2026-07-19: Sekirei S06E04
|
||||
# sat chronic at count 4 despite hasFile=true, already fully resolved) — "consecutive
|
||||
# failures" is supposed to mean consecutive since it last wasn't a problem, not a
|
||||
# cumulative count for all time. Age-skipped items are still in seen_media_ids (added
|
||||
# before the age check above), so a too-new item correctly keeps its count instead of
|
||||
# being reset just for not having been acted on yet.
|
||||
local _fc_key _fc_id
|
||||
for _fc_key in "${!FAILURE_COUNTS[@]}"; do
|
||||
[[ "$_fc_key" == "${arr_type}:"* ]] || continue
|
||||
_fc_id="${_fc_key#${arr_type}:}"
|
||||
[[ -z "${seen_media_ids[$_fc_id]:-}" ]] && unset "FAILURE_COUNTS[$_fc_key]"
|
||||
done
|
||||
|
||||
if [[ "$actioned" -gt 0 ]]; then
|
||||
warn "$arr_name — actioned: $actioned (smart-imported: $smart_imported) | skipped (too new): $skipped_new | chronic: $chronic"
|
||||
else
|
||||
log "$arr_name — nothing actioned | skipped (too new): $skipped_new"
|
||||
fi
|
||||
|
||||
ARR_SUMMARIES+=("$arr_name: actioned $actioned (smart-imported $smart_imported) | too new $skipped_new | chronic $chronic")
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Process Each Arr ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Arrs Failed/Stalled Recovery — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
echo "$ICON_HOST $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
log "Age threshold: ${ARR_IMPORT_RECOVERY_AGE}hr"
|
||||
|
||||
START=$(date +%s)
|
||||
|
||||
# Sonarr — uses aliased vars set by detect_hosts()
|
||||
process_arr \
|
||||
"Sonarr" \
|
||||
"sonarr" \
|
||||
"${SONARR_URL:-}" \
|
||||
"${SONARR_API_KEY:-}" \
|
||||
"v3" \
|
||||
"${SONARR_RECOVERY:-true}" \
|
||||
"${SONARR_VERSION_MAJOR:-4}" \
|
||||
"v3"
|
||||
|
||||
# Radarr — uses aliased vars set by detect_hosts()
|
||||
process_arr \
|
||||
"Radarr" \
|
||||
"radarr" \
|
||||
"${RADARR_URL:-}" \
|
||||
"${RADARR_API_KEY:-}" \
|
||||
"v3" \
|
||||
"${RADARR_RECOVERY:-true}" \
|
||||
"${RADARR_VERSION_MAJOR:-6}" \
|
||||
"v3"
|
||||
|
||||
# Lidarr — HOST1 only, LIDARR_URL empty on HOST2 → exits cleanly via "not configured" guard
|
||||
process_arr \
|
||||
"Lidarr" \
|
||||
"lidarr" \
|
||||
"${LIDARR_URL:-}" \
|
||||
"${LIDARR_API_KEY:-}" \
|
||||
"v1" \
|
||||
"${LIDARR_RECOVERY:-false}" \
|
||||
"${LIDARR_VERSION_MAJOR:-3}" \
|
||||
"v1"
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# Persist updated failure counts — skipped in dry-run so nothing is recorded for a preview
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
mkdir -p "$(dirname "$ARR_RECOVERY_FAILURE_COUNTS")" 2>/dev/null
|
||||
: > "$ARR_RECOVERY_FAILURE_COUNTS"
|
||||
for key in "${!FAILURE_COUNTS[@]}"; do
|
||||
echo "${key}|${FAILURE_COUNTS[$key]}" >> "$ARR_RECOVERY_FAILURE_COUNTS"
|
||||
done
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY ARR RECOVERY SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo "$ICON_TRASH Actioned: $TOTAL_ACTIONED items blocklisted + searched"
|
||||
echo "$ICON_DONE Smart-imported: $TOTAL_SMART_IMPORTED items (upgrade/gap-fill, kept instead of discarded)"
|
||||
echo "$ICON_SKIP Skipped: $TOTAL_SKIPPED items (too new)"
|
||||
echo "$ICON_WARN Chronic: $TOTAL_CHRONIC items (blocklisted, auto re-search stopped)"
|
||||
echo ""
|
||||
for summary in "${ARR_SUMMARIES[@]}"; do
|
||||
echo " $ICON_SUMMARY $summary"
|
||||
done
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ "$TOTAL_ACTIONED" -gt 0 ]]; then
|
||||
warn "$ICON_DONE Done — $TOTAL_ACTIONED item(s) blocklisted and re-searched"
|
||||
notify "Arr recovery on $(hostname) — $TOTAL_ACTIONED item(s) blocklisted and re-searched" \
|
||||
"Arr Recovery" "warning"
|
||||
else
|
||||
echo "$ICON_DONE Done — nothing to recover (all arrs clean)"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# Write stats for sunday_morning_coffee_report.sh
|
||||
if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_RECOVERY_STATS:-}" ]]; then
|
||||
echo "$(date '+%Y-%m-%d')|$(date '+%H:%M')|${TOTAL_ACTIONED}|${TOTAL_SKIPPED}|${TOTAL_CHRONIC}" \
|
||||
>> "$ARR_RECOVERY_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
@@ -8,7 +8,30 @@
|
||||
# Delete orphaned music files not tracked by Lidarr. Queries the API for all
|
||||
# tracked file paths, walks the library on disk, and removes anything untracked
|
||||
# that is old enough to be past the import window. Triggers an Emby library
|
||||
# clean after each deletion run so ghost entries disappear immediately.
|
||||
# clean after runs where files were deleted so ghost entries disappear immediately.
|
||||
#
|
||||
# Pre-flight and the tracked-count floor check are both rescan-aware: a library-wide
|
||||
# rescan legitimately makes tracked counts read low mid-scan (confirmed 2026-07-16 —
|
||||
# 22% of normal during an active RescanFolders), which used to trigger this script's own
|
||||
# hard abort every time it overlapped with a real rescan. Now it checks for an active
|
||||
# rescan-type command first — if one's running, it waits (calibrated to that command's
|
||||
# own historical duration via arr_get_rescan_duration(), up to 3 strikes) and re-fetches
|
||||
# rather than either stacking a duplicate scan or crying wolf on a normal, if slow, state.
|
||||
# Only escalates to the scary abort-and-notify when the count is genuinely low AND nothing
|
||||
# is actively rescanning.
|
||||
#
|
||||
# Cache-first, both layers (2026-07-17). The artist list itself comes from the shared
|
||||
# tracked-data cache via arr_get_tracked_data() — fresh (kept warm every 30min by
|
||||
# arr_cache_prefill.sh), live fetch as fallback. The per-artist trackFile walk below still
|
||||
# always fetches live (that's the actual disk-truth this script's delete decisions depend
|
||||
# on), but write-throughs its result to arr_item_cache_write() so lidarr_missing_art.sh,
|
||||
# running later in the same nightly window, can read it instead of repeating the same walk.
|
||||
# The filesystem is walked once per run, not twice — classification records which paths are
|
||||
# eligible for deletion as it goes, and the delete pass (once the size-threshold check below
|
||||
# passes) just acts on that list instead of re-walking and re-classifying the whole tree.
|
||||
# That single walk also gets size+ctime straight from find -printf instead of a separate stat
|
||||
# fork per file — find already has to stat() every entry to know it's -type f, so this is
|
||||
# free by comparison. Measured ~130x faster per file (0.033ms vs 4.3ms).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
@@ -27,6 +50,28 @@
|
||||
# classification these would be deleted — breaking Lidarr and Emby display.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# API as Ground Truth
|
||||
# What Lidarr tracks is authoritative. Files not in the API response are
|
||||
# orphans — Lidarr has no record of them and they serve no purpose.
|
||||
# The script never infers ownership from directory structure alone.
|
||||
#
|
||||
# Age Gate Before Deletion
|
||||
# Files under LIDARR_ORPHAN_AGE are left alone regardless of tracked status.
|
||||
# Lidarr's import pipeline writes files before registering them — acting
|
||||
# immediately would delete files mid-import.
|
||||
# Age is measured from ctime, not mtime — an import preserves the release's original
|
||||
# mtime, so a file that landed today can read as years old and skip this gate. Depends
|
||||
# on media_shares_permissions.sh touching only entries that are actually wrong.
|
||||
#
|
||||
# Seven-Gate Safety Model
|
||||
# Multiple independent sanity checks must all pass before any file is touched.
|
||||
# No single check is trusted in isolation — a misconfigured path returning an
|
||||
# empty API response must not result in a wiped library.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -41,9 +86,8 @@
|
||||
#
|
||||
# acquire_lock "wait" — large scans take time, wait for previous run to finish
|
||||
# jq + curl validation — exits if either tool missing
|
||||
# DOCKER_TIMEOUT — container checks protected against daemon hangs
|
||||
# ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT)
|
||||
# Duplicate detection — temp file of tracked paths, grep before delete
|
||||
# validate_unraid_cmd — notify script validated before use
|
||||
# Silent by default — orphans/junk warn(), clean library logs silently
|
||||
#
|
||||
# ==============================================================================================
|
||||
@@ -86,7 +130,7 @@
|
||||
# lidarr_cleanup.sh --log — verbose output
|
||||
# lidarr_cleanup.sh --status — show config and exit
|
||||
# lidarr_cleanup.sh --i-know-what-im-doing — bypass size threshold
|
||||
# lidarr_cleanup.sh --i-know-what-im-doing --skip-strike-list — NUCLEAR MODE
|
||||
# lidarr_cleanup.sh --i-know-what-im-doing --skip-age-check — NUCLEAR MODE
|
||||
#
|
||||
# NUCLEAR MODE: both flags bypass age check AND size threshold. Use when Soularr
|
||||
# has filled the gaps and you want a clean one-pass wipe. User accepts full
|
||||
@@ -99,39 +143,14 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# ── Special flag pre-processing ───────────────────────────────────────────────────────────────
|
||||
# Filter --i-know-what-im-doing and --skip-strike-list before parse_args
|
||||
# to avoid unknown flag errors — these are handled separately below.
|
||||
I_KNOW=false
|
||||
SKIP_STRIKES=false
|
||||
FILTERED_ARGS=()
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--i-know-what-im-doing) I_KNOW=true ;;
|
||||
--skip-strike-list) SKIP_STRIKES=true ;;
|
||||
*) FILTERED_ARGS+=("$arg") ;;
|
||||
esac
|
||||
done
|
||||
# Filter --i-know-what-im-doing and --skip-age-check before parse_args
|
||||
# to avoid unknown flag errors — both are handled separately below.
|
||||
parse_destructive_flags "$@"
|
||||
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
|
||||
# ── Nuclear mode warning ──────────────────────────────────────────────────────────────────────
|
||||
if [[ "$I_KNOW" == true ]] && [[ "$SKIP_STRIKES" == true ]] && [[ "$DRY_RUN" != true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo "⚠️ WARNING — NUCLEAR MODE ACTIVE"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo " Flags: --i-know-what-im-doing --skip-strike-list"
|
||||
echo " Strike system: BYPASSED — deletes on first pass"
|
||||
echo " Size threshold: BYPASSED — no GB limit"
|
||||
echo " Data recovery: NOT POSSIBLE after deletion"
|
||||
echo ""
|
||||
echo " Review --dry-run output before proceeding."
|
||||
echo " You have 10 seconds to cancel (Ctrl+C)..."
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
sleep 10
|
||||
echo " Proceeding..."
|
||||
echo ""
|
||||
fi
|
||||
nuclear_mode_warning
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
@@ -153,14 +172,13 @@ if ! command -v jq >/dev/null 2>&1; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
# Lock before detect_hosts — large library scans take time, wait mode appropriate
|
||||
[[ -n "${LIDARR_LOCK_WARN_AGE:-}" ]] && LOCK_WARN_AGE="$LIDARR_LOCK_WARN_AGE"
|
||||
acquire_lock "wait"
|
||||
TMP_DIR="/tmp/lidarr_cleanup_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "_release_all_locks; rm -rf $TMP_DIR" EXIT
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
@@ -176,14 +194,11 @@ if [[ -z "${LIDARR_URL:-}" ]] || [[ -z "${LIDARR_API_KEY:-}" ]]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
DOCKER_TIMEOUT=15
|
||||
ARR_DOCKER_TIMEOUT=15
|
||||
LIDARR_CONTAINER="Lidarr" # container name on HOST1
|
||||
|
||||
# Build path map from HOST1_LIDARR_PATH_MAP for translate_path()
|
||||
declare -A ARR_PATH_MAP
|
||||
for key in "${!HOST1_LIDARR_PATH_MAP[@]}"; do
|
||||
ARR_PATH_MAP["$key"]="${HOST1_LIDARR_PATH_MAP[$key]}"
|
||||
done
|
||||
# Build path map from MY_ID's Lidarr path map
|
||||
build_arr_path_map "LIDARR"
|
||||
|
||||
# Validate required vars — detect_hosts() should have set these
|
||||
require_var LIDARR_URL
|
||||
@@ -197,11 +212,12 @@ if [[ ! -d "$LIDARR_MUSIC_ROOT" ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Config: url=${LIDARR_URL} root=${LIDARR_MUSIC_ROOT}"
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $LIDARR_URL"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be deleted"
|
||||
[[ "$I_KNOW" == true ]] && warn "OVERRIDE — --i-know-what-im-doing active"
|
||||
[[ "$SKIP_STRIKES" == true ]] && warn "OVERRIDE — --skip-strike-list active — age check bypassed"
|
||||
[[ "$SKIP_AGE_CHECK" == true ]] && warn "OVERRIDE — --skip-age-check active — age check bypassed"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
@@ -220,7 +236,7 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo "$ICON_GEAR Protected patterns: ${LIDARR_PROTECTED_PATTERNS[*]}"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "$ICON_GEAR I know: $I_KNOW"
|
||||
echo "$ICON_GEAR Skip strikes: $SKIP_STRIKES"
|
||||
echo "$ICON_GEAR Skip age check: $SKIP_AGE_CHECK"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
@@ -231,83 +247,13 @@ fi
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
|
||||
CONTAINER_RUNNING=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Running}}' "$LIDARR_CONTAINER" 2>/dev/null)
|
||||
if [[ "$CONTAINER_RUNNING" != "true" ]]; then
|
||||
error "$LIDARR_CONTAINER is not running — aborting"
|
||||
notify "Lidarr cleanup aborted on $(hostname) — container not running" \
|
||||
"Lidarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
CONTAINER_HEALTH=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Health.Status}}' "$LIDARR_CONTAINER" 2>/dev/null)
|
||||
case "$CONTAINER_HEALTH" in
|
||||
healthy) info "$LIDARR_CONTAINER is healthy" ;;
|
||||
"") info "$LIDARR_CONTAINER has no health check — proceeding" ;;
|
||||
starting)
|
||||
error "$LIDARR_CONTAINER is still starting — aborting"
|
||||
notify "Lidarr cleanup aborted on $(hostname) — container still starting" \
|
||||
"Lidarr Cleanup" "warning"
|
||||
exit 1 ;;
|
||||
unhealthy)
|
||||
error "$LIDARR_CONTAINER is unhealthy — aborting"
|
||||
notify "Lidarr cleanup aborted on $(hostname) — container unhealthy" \
|
||||
"Lidarr Cleanup" "warning"
|
||||
exit 1 ;;
|
||||
*) warn "$LIDARR_CONTAINER health: $CONTAINER_HEALTH — proceeding with caution" ;;
|
||||
esac
|
||||
|
||||
info "Safety layer 1 passed — container healthy"
|
||||
check_container_health "$LIDARR_CONTAINER" "$ARR_DOCKER_TIMEOUT" "Lidarr Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Lidarr API call with HTTP status check
|
||||
# Usage: lidarr_api "artist" | lidarr_api "trackFile?artistId=123"
|
||||
lidarr_api() {
|
||||
local endpoint="$1"
|
||||
local response http_code body
|
||||
|
||||
response=$(curl -sf \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${LIDARR_URL}/api/v1/${endpoint}" 2>/dev/null)
|
||||
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
|
||||
if [[ "$http_code" != "200" ]]; then
|
||||
error "Lidarr API HTTP $http_code for: $endpoint"
|
||||
return 1
|
||||
fi
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
# Check if a file extension is a tracked music format
|
||||
is_music_file() {
|
||||
local ext="${1##*.}"
|
||||
ext="${ext,,}"
|
||||
for valid_ext in "${LIDARR_EXTENSIONS[@]}"; do
|
||||
[[ "$ext" == "$valid_ext" ]] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
# Check if a file matches any protected pattern
|
||||
is_protected_file() {
|
||||
local filename
|
||||
filename=$(basename "$1")
|
||||
for pattern in "${LIDARR_PROTECTED_PATTERNS[@]}"; do
|
||||
# shellcheck disable=SC2254
|
||||
case "$filename" in
|
||||
$pattern) return 0 ;;
|
||||
esac
|
||||
done
|
||||
return 1
|
||||
}
|
||||
# check_container_health(), arr_api(), has_extension(), matches_pattern_list() — common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Pre-flight: Lidarr Import Scan ━━━
|
||||
@@ -325,43 +271,30 @@ for _cp in "${!ARR_PATH_MAP[@]}"; do
|
||||
done
|
||||
unset _cp
|
||||
|
||||
if [[ -n "$LIDARR_CONTAINER_ROOT" ]]; then
|
||||
# Don't stack a fresh scan on top of one already running — confirmed 2026-07-16 that
|
||||
# repeated runs each firing their own DownloadedAlbumsScan piled up in Lidarr's command
|
||||
# queue behind each other rather than replacing/coalescing, contributing to a multi-hour
|
||||
# backlog. If something's already scanning, just wait for that one instead.
|
||||
_already_active=$(arr_active_rescan_command "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")
|
||||
|
||||
if [[ -n "$_already_active" ]]; then
|
||||
info "$_already_active already in progress — waiting for it instead of starting a new scan"
|
||||
_wait=$(( $(arr_get_rescan_duration "lidarr" "$_already_active" "${LIDARR_IMPORT_SCAN_TIMEOUT:-600}") ))
|
||||
_polled=0
|
||||
while [[ "$_polled" -lt "$_wait" ]]; do
|
||||
[[ -z "$(arr_active_rescan_command "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")" ]] && break
|
||||
sleep 15
|
||||
(( _polled += 15 ))
|
||||
[[ $(( _polled % 60 )) -eq 0 ]] && log " Still waiting on $_already_active... (${_polled}s elapsed)"
|
||||
done
|
||||
elif [[ -n "$LIDARR_CONTAINER_ROOT" ]]; then
|
||||
info "Triggering DownloadedAlbumsScan on: $LIDARR_CONTAINER_ROOT"
|
||||
SCAN_PAYLOAD="{\"name\": \"DownloadedAlbumsScan\", \"path\": \"$LIDARR_CONTAINER_ROOT\"}"
|
||||
trigger_and_await_command "$LIDARR_URL" "$LIDARR_API_KEY" "v1" "$SCAN_PAYLOAD" "${LIDARR_IMPORT_SCAN_TIMEOUT:-600}" "lidarr"
|
||||
else
|
||||
info "No path map match — triggering DownloadedAlbumsScan (all root folders)"
|
||||
SCAN_PAYLOAD='{"name": "DownloadedAlbumsScan"}'
|
||||
fi
|
||||
|
||||
SCAN_RESPONSE=$(curl -sf --max-time 30 -X POST \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$SCAN_PAYLOAD" \
|
||||
"${LIDARR_URL}/api/v1/command" 2>/dev/null)
|
||||
|
||||
SCAN_CMD_ID=$(echo "$SCAN_RESPONSE" | jq -r '.id // empty' 2>/dev/null)
|
||||
|
||||
if [[ -z "$SCAN_CMD_ID" ]]; then
|
||||
warn "Could not trigger import scan — proceeding without pre-flight"
|
||||
else
|
||||
info "Import scan queued (command ID: $SCAN_CMD_ID) — waiting for completion..."
|
||||
POLL_TIMEOUT=${LIDARR_IMPORT_SCAN_TIMEOUT:-600}
|
||||
POLLED=0
|
||||
while [[ "$POLLED" -lt "$POLL_TIMEOUT" ]]; do
|
||||
SCAN_STATUS=$(curl -sf --max-time 10 \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" \
|
||||
"${LIDARR_URL}/api/v1/command/${SCAN_CMD_ID}" 2>/dev/null | \
|
||||
jq -r '.status // empty' 2>/dev/null)
|
||||
case "$SCAN_STATUS" in
|
||||
completed) info "Import scan complete ✅"; break ;;
|
||||
failed) warn "Import scan reported failed — proceeding anyway"; break ;;
|
||||
esac
|
||||
sleep 10
|
||||
(( POLLED += 10 ))
|
||||
[[ $(( POLLED % 60 )) -eq 0 ]] && log " Still scanning... (${POLLED}s elapsed)"
|
||||
done
|
||||
[[ "$POLLED" -ge "$POLL_TIMEOUT" ]] && \
|
||||
warn "Import scan timed out after ${POLL_TIMEOUT}s — proceeding anyway"
|
||||
trigger_and_await_command "$LIDARR_URL" "$LIDARR_API_KEY" "v1" "$SCAN_PAYLOAD" "${LIDARR_IMPORT_SCAN_TIMEOUT:-600}" "lidarr"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -381,8 +314,12 @@ check_arr_version "$LIDARR_URL" "$LIDARR_API_KEY" "v1" "$LIDARR_VERSION_MAJOR" "
|
||||
|
||||
info "Querying Lidarr API..."
|
||||
|
||||
# Fetch all artists
|
||||
ARTIST_RESPONSE=$(lidarr_api "artist") || {
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (now kept
|
||||
# current every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to
|
||||
# a live fetch when it's stale, and waits out an active rescan before either. Only the artist
|
||||
# list itself is cached — the per-artist trackFile data below is never cached and always live,
|
||||
# since that's the actual disk-truth this script's cleanup decisions depend on.
|
||||
ARTIST_RESPONSE=$(arr_get_tracked_data "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1") || {
|
||||
error "Failed to fetch artists from Lidarr"
|
||||
notify "Lidarr cleanup failed on $(hostname) — could not fetch artists" \
|
||||
"Lidarr Cleanup" "warning"
|
||||
@@ -390,7 +327,7 @@ ARTIST_RESPONSE=$(lidarr_api "artist") || {
|
||||
}
|
||||
|
||||
ARTIST_IDS=$(echo "$ARTIST_RESPONSE" | jq -r '.[].id' 2>/dev/null)
|
||||
ARTIST_COUNT=$(echo "$ARTIST_IDS" | grep -c "[0-9]" 2>/dev/null || echo 0)
|
||||
ARTIST_COUNT=$(echo "$ARTIST_IDS" | grep -c "[0-9]" 2>/dev/null || true)
|
||||
|
||||
# Safety Layer 4 — artist count > 0
|
||||
if [[ "$ARTIST_COUNT" -eq 0 ]]; then
|
||||
@@ -402,35 +339,50 @@ fi
|
||||
|
||||
info "Found $ARTIST_COUNT artists — fetching track files..."
|
||||
|
||||
TMP_DIR="/tmp/lidarr_cleanup_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "rm -rf $TMP_DIR" EXIT
|
||||
|
||||
TRACKED_FILE="$TMP_DIR/tracked_paths.txt"
|
||||
> "$TRACKED_FILE"
|
||||
|
||||
while IFS= read -r artist_id; do
|
||||
[[ -z "$artist_id" ]] && continue
|
||||
ARTIST_TRACKS=$(lidarr_api "trackFile?artistId=${artist_id}" 2>/dev/null)
|
||||
if [[ -n "$ARTIST_TRACKS" ]]; then
|
||||
while IFS= read -r api_path; do
|
||||
[[ -z "$api_path" ]] && continue
|
||||
translate_path "$api_path" >> "$TRACKED_FILE"
|
||||
done < <(echo "$ARTIST_TRACKS" | jq -r '.[].path' 2>/dev/null)
|
||||
fi
|
||||
done <<< "$ARTIST_IDS"
|
||||
# Fetches every artist's track-file paths fresh into TRACKED_FILE/TRACKED_MAP/TRACKED_COUNT.
|
||||
# Pulled into a function so the rescan-aware retry below can re-fetch after waiting without
|
||||
# duplicating this whole loop inline.
|
||||
#
|
||||
# Write-through — also caches the raw per-track data via arr_item_cache_write() (2026-07-17)
|
||||
# so scripts running later in the same maintenance window (lidarr_missing_art.sh) can read it
|
||||
# instead of repeating this same per-artist walk. This walk is happening regardless for our
|
||||
# own cleanup decisions; the cache write is free by comparison. See common.sh for the pattern.
|
||||
_fetch_tracked_files() {
|
||||
> "$TRACKED_FILE"
|
||||
local all_tracks_tmp
|
||||
all_tracks_tmp=$(mktemp)
|
||||
while IFS= read -r artist_id; do
|
||||
[[ -z "$artist_id" ]] && continue
|
||||
ARTIST_TRACKS=$(arr_api "$LIDARR_URL" "$LIDARR_API_KEY" "v1" "trackFile?artistId=${artist_id}" "Lidarr" 2>/dev/null)
|
||||
if [[ -n "$ARTIST_TRACKS" ]]; then
|
||||
echo "$ARTIST_TRACKS" >> "$all_tracks_tmp"
|
||||
while IFS= read -r api_path; do
|
||||
[[ -z "$api_path" ]] && continue
|
||||
translate_path "$api_path" >> "$TRACKED_FILE"
|
||||
done < <(echo "$ARTIST_TRACKS" | jq -r '.[].path' 2>/dev/null)
|
||||
fi
|
||||
done <<< "$ARTIST_IDS"
|
||||
|
||||
sort -u "$TRACKED_FILE" -o "$TRACKED_FILE"
|
||||
arr_item_cache_write "lidarr" "$(jq -s 'add // []' "$all_tracks_tmp" 2>/dev/null)"
|
||||
rm -f "$all_tracks_tmp"
|
||||
|
||||
# Build in-memory lookup map — O(1) per lookup vs O(n) grep per file
|
||||
# Eliminates the main performance bottleneck for large libraries
|
||||
declare -A TRACKED_MAP
|
||||
while IFS= read -r _tracked_path; do
|
||||
[[ -n "$_tracked_path" ]] && TRACKED_MAP["$_tracked_path"]=1
|
||||
done < "$TRACKED_FILE"
|
||||
unset _tracked_path
|
||||
sort -u "$TRACKED_FILE" -o "$TRACKED_FILE"
|
||||
|
||||
# Build in-memory lookup map — O(1) per lookup vs O(n) grep per file
|
||||
# Eliminates the main performance bottleneck for large libraries
|
||||
unset TRACKED_MAP
|
||||
declare -gA TRACKED_MAP
|
||||
while IFS= read -r _tracked_path; do
|
||||
[[ -n "$_tracked_path" ]] && TRACKED_MAP["$_tracked_path"]=1
|
||||
done < "$TRACKED_FILE"
|
||||
unset _tracked_path
|
||||
TRACKED_COUNT=$(wc -l < "$TRACKED_FILE")
|
||||
}
|
||||
|
||||
_fetch_tracked_files
|
||||
info "Built in-memory lookup map: ${#TRACKED_MAP[@]} tracked paths"
|
||||
TRACKED_COUNT=$(wc -l < "$TRACKED_FILE")
|
||||
|
||||
# Safety Layer 5 — tracked count > 0
|
||||
if [[ "$TRACKED_COUNT" -eq 0 ]]; then
|
||||
@@ -442,26 +394,42 @@ fi
|
||||
|
||||
info "$ARTIST_COUNT artists | $TRACKED_COUNT tracked files"
|
||||
|
||||
# Safety Layer 6 — percentage drop vs last known count
|
||||
if [[ -f "$LIDARR_TRACKED_COUNT_FILE" ]]; then
|
||||
LAST_COUNT=$(cat "$LIDARR_TRACKED_COUNT_FILE" 2>/dev/null || echo 0)
|
||||
if [[ "$LAST_COUNT" -gt 0 ]]; then
|
||||
PCT=$(awk "BEGIN {printf \"%d\", ($TRACKED_COUNT / $LAST_COUNT) * 100}")
|
||||
if [[ "$PCT" -lt "$LIDARR_MIN_TRACKED_PCT" ]]; then
|
||||
error "Tracked count dropped to ${PCT}% of last run ($TRACKED_COUNT vs $LAST_COUNT)"
|
||||
error "Suggests API issue — aborting to prevent mass deletion"
|
||||
error "If expected (large library removal) delete: $LIDARR_TRACKED_COUNT_FILE"
|
||||
notify "Lidarr cleanup aborted on $(hostname) — tracked count dropped to ${PCT}%" \
|
||||
"Lidarr Cleanup" "warning"
|
||||
exit 1
|
||||
# Safety Layer 6 — percentage drop vs last known count, with rescan-aware retry.
|
||||
# A library-wide rescan legitimately makes tracked counts read low mid-scan — sometimes
|
||||
# dramatically (confirmed 2026-07-16: 22% of normal during an active RescanFolders).
|
||||
# That's not "something's wrong," it's Lidarr actively re-verifying every file. Wait it out
|
||||
# (calibrated to that command's own historical duration) before treating a drop as a genuine
|
||||
# problem worth the scary abort-and-notify. Only escalates to the hard abort in
|
||||
# check_tracked_count_floor if the count is still low AND nothing is actively rescanning —
|
||||
# that combination is the actually-suspicious case the floor check exists to catch.
|
||||
_last_known=$(cat "$LIDARR_TRACKED_COUNT_FILE" 2>/dev/null || echo 0)
|
||||
if [[ "$_last_known" -gt 0 ]]; then
|
||||
_strike=1
|
||||
while [[ "$_strike" -le 3 ]]; do
|
||||
_pct=$(awk "BEGIN {printf \"%d\", ($TRACKED_COUNT / $_last_known) * 100}")
|
||||
[[ "$_pct" -ge "${LIDARR_MIN_TRACKED_PCT:-50}" ]] && break
|
||||
|
||||
_active_cmd=$(arr_active_rescan_command "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")
|
||||
[[ -z "$_active_cmd" ]] && break # low count, nothing rescanning — genuine, don't retry
|
||||
|
||||
_wait=$(( $(arr_get_rescan_duration "lidarr" "$_active_cmd" 300) / 2 ))
|
||||
[[ "$_wait" -lt 30 ]] && _wait=30
|
||||
warn "Tracked count ${_pct}% of last run, but $_active_cmd active — waiting ${_wait}s (strike ${_strike}/3)"
|
||||
sleep "$_wait"
|
||||
_fetch_tracked_files
|
||||
(( _strike++ ))
|
||||
done
|
||||
|
||||
if [[ "$_strike" -gt 3 ]]; then
|
||||
_active_cmd=$(arr_active_rescan_command "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")
|
||||
if [[ -n "$_active_cmd" ]]; then
|
||||
warn "Lidarr still busy ($_active_cmd) after 3 strikes — deferring to next scheduled run"
|
||||
exit 0
|
||||
fi
|
||||
info "Tracked count: ${PCT}% of last run ($TRACKED_COUNT vs $LAST_COUNT) ✅"
|
||||
fi
|
||||
else
|
||||
info "No previous count on record — first run, saving baseline"
|
||||
fi
|
||||
|
||||
echo "$TRACKED_COUNT" > "$LIDARR_TRACKED_COUNT_FILE"
|
||||
check_tracked_count_floor "$TRACKED_COUNT" "$LIDARR_TRACKED_COUNT_FILE" "$LIDARR_MIN_TRACKED_PCT" "Lidarr Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Scan Music Root ━━━
|
||||
@@ -480,10 +448,42 @@ JUNK_BYTES=0
|
||||
|
||||
AGE_SECONDS=$(( LIDARR_ORPHAN_AGE * 86400 ))
|
||||
NOW=$(date +%s)
|
||||
MAX_DELETE_BYTES=$(awk "BEGIN {printf \"%d\", $LIDARR_MAX_DELETE_GB * 1073741824}")
|
||||
|
||||
while IFS= read -r filepath; do
|
||||
# Files classified ORPHAN/JUNK below get their path recorded here, so the deletion pass can
|
||||
# just delete them directly instead of re-walking and re-classifying the whole tree a second
|
||||
# time (2026-07-17) — the size-threshold check below needs to know the total before deleting
|
||||
# anything, not before knowing what to delete, so there's no need to redo the classification
|
||||
# itself once it's already been decided.
|
||||
TO_DELETE_FILE="$TMP_DIR/to_delete_paths.txt"
|
||||
> "$TO_DELETE_FILE"
|
||||
|
||||
# ── Orphan strikes ────────────────────────────────────────────────────────────────────────────
|
||||
# Same contract as radarr_cleanup.sh: a file must classify for deletion on
|
||||
# LIDARR_ORPHAN_STRIKE_LIMIT consecutive runs before it is removed. Covers the partial
|
||||
# classification failure that is too small to trip the tracked-count floor above. The file is
|
||||
# rebuilt from each run rather than edited, which is what prunes it.
|
||||
LIDARR_ORPHAN_STRIKE_LIMIT="${LIDARR_ORPHAN_STRIKE_LIMIT:-2}"
|
||||
STRIKES_FILE="${LIDARR_ORPHAN_STRIKES_FILE:-$DB_DIR/lidarr_orphan_strikes.tsv}"
|
||||
mkdir -p "$(dirname "$STRIKES_FILE")" 2>/dev/null || true
|
||||
touch "$STRIKES_FILE" 2>/dev/null || true
|
||||
STRIKES_NEW="$TMP_DIR/strikes_new.tsv"
|
||||
> "$STRIKES_NEW"
|
||||
HELD_COUNT=0
|
||||
HELD_BYTES=0
|
||||
|
||||
orphan_strike_ok() {
|
||||
local path="$1" prev strikes
|
||||
prev=$(wd_state_get "$path" "$STRIKES_FILE"); prev="${prev//[^0-9]/}"
|
||||
strikes=$(( ${prev:-0} + 1 ))
|
||||
printf '%s:%s\n' "$path" "$strikes" >> "$STRIKES_NEW"
|
||||
(( strikes >= LIDARR_ORPHAN_STRIKE_LIMIT )) && return 0
|
||||
warn " strike $strikes/$LIDARR_ORPHAN_STRIKE_LIMIT — not removing yet: $path"
|
||||
return 1
|
||||
}
|
||||
|
||||
while read -r FILE_SIZE FILE_CTIME filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
FILE_CTIME="${FILE_CTIME%%.*}"
|
||||
|
||||
# Tracked — leave alone
|
||||
if [[ -n "${TRACKED_MAP[$filepath]:-}" ]]; then
|
||||
@@ -492,19 +492,23 @@ while IFS= read -r filepath; do
|
||||
fi
|
||||
|
||||
# Protected — never delete
|
||||
if is_protected_file "$filepath"; then
|
||||
if matches_pattern_list "$filepath" "${LIDARR_PROTECTED_PATTERNS[@]}"; then
|
||||
log "$ICON_PROTECTED PROTECTED: $filepath"
|
||||
(( PROTECTED_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
FILE_SIZE=$(stat -c%s "$filepath" 2>/dev/null || echo 0)
|
||||
if has_extension "$filepath" "${LIDARR_EXTENSIONS[@]}"; then
|
||||
# ctime, not mtime — an import preserves the release's original mtime, so a file
|
||||
# Lidarr moved in today can read as years old and skip this gate entirely.
|
||||
# Measured 2026-07-27: 400 of 400 files imported that week had mtimes over 7
|
||||
# days, one of them 9613 days. ctime is stamped when the file lands on this
|
||||
# filesystem and cannot be carried in from an archive. This only holds because
|
||||
# media_shares_permissions.sh applies owner/mode conditionally — a blanket
|
||||
# chown/chmod restamps every inode nightly and would peg every file at age 0.
|
||||
FILE_AGE=$(( NOW - FILE_CTIME ))
|
||||
|
||||
if is_music_file "$filepath"; then
|
||||
FILE_MTIME=$(stat -c %Y "$filepath" 2>/dev/null || echo 0)
|
||||
FILE_AGE=$(( NOW - FILE_MTIME ))
|
||||
|
||||
if [[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && [[ "$SKIP_STRIKES" != true ]]; then
|
||||
if [[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && [[ "$SKIP_AGE_CHECK" != true ]]; then
|
||||
log "RECENT (skipping): $filepath"
|
||||
(( RECENT_COUNT++ ))
|
||||
continue
|
||||
@@ -513,55 +517,67 @@ while IFS= read -r filepath; do
|
||||
warn "$ICON_TRASH ORPHAN: $filepath"
|
||||
(( ORPHAN_COUNT++ ))
|
||||
ORPHAN_BYTES=$(( ORPHAN_BYTES + FILE_SIZE ))
|
||||
if ! orphan_strike_ok "$filepath"; then (( HELD_COUNT++ )); HELD_BYTES=$(( HELD_BYTES + FILE_SIZE )); continue; fi
|
||||
printf '%s\t%s\t%s\n' "$FILE_SIZE" "$FILE_CTIME" "$filepath" >> "$TO_DELETE_FILE"
|
||||
else
|
||||
log "JUNK: $filepath"
|
||||
(( JUNK_COUNT++ ))
|
||||
JUNK_BYTES=$(( JUNK_BYTES + FILE_SIZE ))
|
||||
if ! orphan_strike_ok "$filepath"; then (( HELD_COUNT++ )); HELD_BYTES=$(( HELD_BYTES + FILE_SIZE )); continue; fi
|
||||
printf '%s\t%s\t%s\n' "$FILE_SIZE" "$FILE_CTIME" "$filepath" >> "$TO_DELETE_FILE"
|
||||
fi
|
||||
|
||||
done < <(find "$LIDARR_MUSIC_ROOT" -type f 2>/dev/null)
|
||||
# -printf gets size + mtime directly from find's own stat() during the walk, instead of a
|
||||
# separate stat fork per file (2026-07-17) — measured ~130x faster per file (0.033ms vs
|
||||
# 4.3ms), since find already has to stat() every entry anyway to know it's -type f.
|
||||
done < <(find "$LIDARR_MUSIC_ROOT" -type f -printf '%s %C@ %p\n' 2>/dev/null)
|
||||
|
||||
TOTAL_DELETE_BYTES=$(( ORPHAN_BYTES + JUNK_BYTES ))
|
||||
TOTAL_REMOVED=$(( ORPHAN_COUNT + JUNK_COUNT ))
|
||||
# Eligible, not classified: a file still serving its strikes is an orphan but is not queued this
|
||||
# run, so it must not appear in the denominator the budget reports against.
|
||||
TOTAL_DELETE_BYTES=$(( ORPHAN_BYTES + JUNK_BYTES - HELD_BYTES ))
|
||||
TOTAL_REMOVED=$(( ORPHAN_COUNT + JUNK_COUNT - HELD_COUNT ))
|
||||
|
||||
# Rebuilt, never edited. Skipped on a dry run: a preview that advanced real counters would make
|
||||
# the next real run delete a run early.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
mv "$STRIKES_NEW" "$STRIKES_FILE" 2>/dev/null || warn "Could not update $STRIKES_FILE"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 7 — Deletion Size Threshold ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$TOTAL_DELETE_BYTES" -gt "$MAX_DELETE_BYTES" ]]; then
|
||||
TOTAL_HUMAN=$(awk "BEGIN {printf \"%.1fGB\", $TOTAL_DELETE_BYTES / 1073741824}")
|
||||
if [[ "$I_KNOW" != true ]]; then
|
||||
echo ""
|
||||
error "Deletion would exceed ${LIDARR_MAX_DELETE_GB}GB — $TOTAL_HUMAN would be deleted"
|
||||
error "Review ORPHAN lines above carefully before proceeding"
|
||||
error "Rerun with: --i-know-what-im-doing"
|
||||
error "To also bypass age check: add --skip-strike-list"
|
||||
notify "Lidarr cleanup halted on $(hostname) — ${TOTAL_HUMAN} requires --i-know-what-im-doing" \
|
||||
# A per-run budget, not a veto — see apply_delete_budget() in common.sh. The ceiling still caps
|
||||
# any single run; it just no longer deadlocks on a backlog larger than itself.
|
||||
BUDGET_FILE="$TMP_DIR/to_delete_budgeted.txt"
|
||||
|
||||
if [[ "$I_KNOW" == true ]]; then
|
||||
warn "OVERRIDE — --i-know-what-im-doing active, per-run budget not applied"
|
||||
cut -d"$(printf '\t')" -f3- "$TO_DELETE_FILE" > "$BUDGET_FILE"
|
||||
_BUDGET_KEPT_COUNT=$TOTAL_REMOVED; _BUDGET_KEPT_BYTES=$TOTAL_DELETE_BYTES
|
||||
_BUDGET_DEFERRED_COUNT=0; _BUDGET_DEFERRED_BYTES=0; _BUDGET_STUCK=""
|
||||
else
|
||||
apply_delete_budget "$TO_DELETE_FILE" "$BUDGET_FILE" "$LIDARR_MAX_DELETE_GB"
|
||||
if [[ -n "$_BUDGET_STUCK" ]]; then
|
||||
error "Single file exceeds the ${LIDARR_MAX_DELETE_GB}GB budget on its own — nothing removed this run"
|
||||
error " $_BUDGET_STUCK"
|
||||
error "Raise LIDARR_MAX_DELETE_GB or clear this one with --i-know-what-im-doing"
|
||||
notify "Lidarr cleanup stalled on $(hostname) — one file exceeds the ${LIDARR_MAX_DELETE_GB}GB budget" \
|
||||
"Lidarr Cleanup" "warning"
|
||||
exit 1
|
||||
else
|
||||
warn "OVERRIDE — deletion is $TOTAL_HUMAN — proceeding with --i-know-what-im-doing"
|
||||
elif [[ "$_BUDGET_DEFERRED_COUNT" -gt 0 ]]; then
|
||||
warn "Budget ${LIDARR_MAX_DELETE_GB}GB — removing $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED ($(format_bytes "$_BUDGET_KEPT_BYTES")), deferring $_BUDGET_DEFERRED_COUNT ($(format_bytes "$_BUDGET_DEFERRED_BYTES")) to the next run"
|
||||
notify "Lidarr cleanup removed $(format_bytes "$_BUDGET_KEPT_BYTES") of $(format_bytes "$TOTAL_DELETE_BYTES") on $(hostname) — $_BUDGET_DEFERRED_COUNT file(s) deferred" \
|
||||
"Lidarr Cleanup" "normal"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Execute Deletions ─────────────────────────────────────────────────────────────────────────
|
||||
# All safety layers passed — delete orphans and junk
|
||||
# All safety layers passed — delete orphans and junk. Reuses TO_DELETE_FILE from the
|
||||
# classification pass above instead of re-walking and re-classifying the whole tree again.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
[[ -n "${TRACKED_MAP[$filepath]:-}" ]] && continue
|
||||
is_protected_file "$filepath" && continue
|
||||
|
||||
FILE_MTIME=$(stat -c %Y "$filepath" 2>/dev/null || echo 0)
|
||||
FILE_AGE=$(( NOW - FILE_MTIME ))
|
||||
|
||||
if is_music_file "$filepath"; then
|
||||
[[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && \
|
||||
[[ "$SKIP_STRIKES" != true ]] && continue
|
||||
fi
|
||||
|
||||
rm -f "$filepath" 2>/dev/null || error "Failed to delete: $filepath"
|
||||
|
||||
done < <(find "$LIDARR_MUSIC_ROOT" -type f 2>/dev/null)
|
||||
done < "$BUDGET_FILE"
|
||||
|
||||
info "Cleaning up empty folders..."
|
||||
find "$LIDARR_MUSIC_ROOT" -mindepth 1 -type d -empty -delete 2>/dev/null
|
||||
@@ -570,16 +586,7 @@ fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
format_bytes() {
|
||||
local bytes=$1
|
||||
if (( bytes > 1073741824 )); then
|
||||
awk "BEGIN {printf \"%.1fGB\", $bytes / 1073741824}"
|
||||
elif (( bytes > 1048576 )); then
|
||||
awk "BEGIN {printf \"%.1fMB\", $bytes / 1048576}"
|
||||
else
|
||||
echo "${bytes}B"
|
||||
fi
|
||||
}
|
||||
# format_bytes() — provided by common.sh
|
||||
|
||||
ORPHAN_HUMAN=$(format_bytes "$ORPHAN_BYTES")
|
||||
JUNK_HUMAN=$(format_bytes "$JUNK_BYTES")
|
||||
@@ -595,6 +602,10 @@ echo "$ICON_SHIELD Protected: $PROTECTED_COUNT files (cover art, metadata
|
||||
echo "$ICON_TRASH Orphans: $ORPHAN_COUNT files ($ORPHAN_HUMAN)"
|
||||
echo "$ICON_TRASH Junk: $JUNK_COUNT files ($JUNK_HUMAN)"
|
||||
echo "$ICON_SKIP Recent skipped: $RECENT_COUNT files (under ${LIDARR_ORPHAN_AGE} days)"
|
||||
[[ "${HELD_COUNT:-0}" -gt 0 ]] && \
|
||||
echo "$ICON_SKIP Held (strikes): $HELD_COUNT files ($(format_bytes "$HELD_BYTES")) — under ${LIDARR_ORPHAN_STRIKE_LIMIT} consecutive runs"
|
||||
[[ "${_BUDGET_DEFERRED_COUNT:-0}" -gt 0 ]] && \
|
||||
echo "$ICON_SKIP Deferred: $_BUDGET_DEFERRED_COUNT files ($(format_bytes "$_BUDGET_DEFERRED_BYTES")) — over the ${LIDARR_MAX_DELETE_GB}GB run budget"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
@@ -603,8 +614,10 @@ if [[ "$DRY_RUN" == true ]]; then
|
||||
elif [[ "$TOTAL_REMOVED" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Clean — nothing to remove"
|
||||
else
|
||||
warn "$ICON_DONE Removed $TOTAL_REMOVED files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)"
|
||||
notify "Lidarr cleanup on $(hostname) — removed $TOTAL_REMOVED files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)" "Lidarr Cleanup" "warning"
|
||||
# What was actually removed, not what was classified. With strikes and a budget in force those
|
||||
# differ, and reporting the classification as the outcome is the oldest bug shape here.
|
||||
warn "$ICON_DONE Removed $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED eligible files ($(format_bytes "$_BUDGET_KEPT_BYTES"))"
|
||||
notify "Lidarr cleanup on $(hostname) — removed $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED eligible files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)" "Lidarr Cleanup" "warning"
|
||||
# Notify Emby to clean missing files — removes ghost entries immediately
|
||||
notify_emby_scan
|
||||
fi
|
||||
@@ -616,4 +629,4 @@ if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_CLEANUP_STATS:-}" ]]; then
|
||||
>> "$ARR_CLEANUP_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
exit 0
|
||||
+363
@@ -0,0 +1,363 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ======================= Lidarr Duplicate Artist Cleanup =======================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Detects and resolves duplicate artist entries in Lidarr's library — cases where the
|
||||
# same display name (case-insensitive) is backed by two different MusicBrainz artist
|
||||
# IDs. This happens when a search or list sync matches a same-named-but-different real
|
||||
# artist and adds it alongside the one already in the library. From that point on,
|
||||
# every completed download for that display name throws MultipleArtistsFoundException
|
||||
# and can never import — Lidarr correctly refuses to guess which of the two it means.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# For each case-insensitive duplicate artist name found:
|
||||
#
|
||||
# Only one side has any tracked files
|
||||
# → The zero-file side is a phantom — delete it (deleteFiles=false, nothing on disk
|
||||
# to lose) and add its MusicBrainz ID to Lidarr's Import List Exclusions so it
|
||||
# can't be silently re-added by a future list sync. Then trigger a refresh on the
|
||||
# surviving artist so anything that was stuck on this exact ambiguity resolves
|
||||
# immediately instead of waiting for Lidarr's own next check cycle.
|
||||
#
|
||||
# Both sides have files, but their album titles don't overlap at all
|
||||
# → Genuinely two different real artists sharing a name (e.g. a band's classic
|
||||
# lineup vs. a later solo era with the same stage name). Not a bug — left alone.
|
||||
#
|
||||
# Both sides have files AND overlapping album titles
|
||||
# → The one case actually risky to automate: could mean real content is split
|
||||
# across two entries and needs an actual merge, not a delete. Untouched,
|
||||
# notified for manual review.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Never Delete Real Content
|
||||
# Only the zero-tracked-file side of a pair is ever deleted. Anything with files is
|
||||
# either left alone (disjoint albums) or flagged for a human (overlapping albums) —
|
||||
# never auto-removed.
|
||||
#
|
||||
# Block Re-Addition At The Source
|
||||
# A phantom that keeps coming back is worse than one that was never cleaned —
|
||||
# Import List Exclusion is Lidarr's own mechanism for "never auto-add this again."
|
||||
#
|
||||
# Silent When Clean
|
||||
# No duplicates found, or all duplicates are the simple phantom case — minimal
|
||||
# output. Only ambiguous (overlapping-album) pairs produce a notification.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Required by the container interaction and state writes.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents overlapping runs racing on the same artist IDs.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases LIDARR_URL / LIDARR_API_KEY.
|
||||
#
|
||||
# jq Dependency Check
|
||||
# Fails fast if jq is missing — duplicate detection and every file-count read
|
||||
# depend on it, and an absent jq would evaluate counts to empty and make every
|
||||
# artist look like a zero-file phantom.
|
||||
#
|
||||
# API Reachability + Version Gate
|
||||
# check_api then check_arr_version against LIDARR_VERSION_MAJOR before any read.
|
||||
#
|
||||
# Empty Library Abort
|
||||
# A response of 0 artists aborts. An empty list is indistinguishable from
|
||||
# "no duplicates" and must never be read as a clean result.
|
||||
#
|
||||
# Tracked-Count Floor
|
||||
# check_tracked_count_floor against the baseline shared with lidarr_cleanup.sh.
|
||||
# A library-wide desync — mid full-rescan, for example — makes trackFileCount read
|
||||
# far below reality for many artists at once. Confirmed 2026-07-16: without this,
|
||||
# both sides of a genuinely-real duplicate (ROMES) read as 0-file phantoms and the
|
||||
# wrong one would have been deleted. Deliberately reuses lidarr_cleanup.sh's own
|
||||
# baseline so every script depending on tracked counts shares one answer to "is
|
||||
# Lidarr's data trustworthy right now" rather than forming a separate opinion.
|
||||
#
|
||||
# Files Never Deleted
|
||||
# Removal passes deleteFiles=false. Only the phantom Lidarr entry is dropped;
|
||||
# nothing on disk is touched, so a wrong call costs a re-add, not media.
|
||||
#
|
||||
# Phantom-Only Deletion
|
||||
# Only the zero-file side of a duplicate pair is ever removed. If both sides hold
|
||||
# files, or neither does unambiguously, the pair is flagged for manual review
|
||||
# instead — the script never picks a winner between two real artists.
|
||||
#
|
||||
# Manual Review Reporting
|
||||
# Flagged pairs are named in the summary and notification so an ambiguous
|
||||
# duplicate surfaces as a decision to make rather than disappearing silently.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run reports every deletion and exclusion and performs none.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY — aliased by detect_hosts()
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# lidarr_duplicate_artist_cleanup.sh — normal run
|
||||
# lidarr_duplicate_artist_cleanup.sh --dry-run — preview, no deletions or exclusions
|
||||
# lidarr_duplicate_artist_cleanup.sh --log — verbose per-pair output
|
||||
# lidarr_duplicate_artist_cleanup.sh --status — show config and exit
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
notify "lidarr_duplicate_artist_cleanup failed on $(hostname) — jq not installed" \
|
||||
"Lidarr Duplicate Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock "wait"
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases LIDARR_URL, LIDARR_API_KEY
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "${LIDARR_URL:-}" ]] || [[ -z "${LIDARR_API_KEY:-}" ]]; then
|
||||
info "Lidarr not configured on $MY_ID ($LOCAL_SERVER_NAME) — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no artists will be deleted or excluded"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Lidarr URL: $LIDARR_URL"
|
||||
echo "$ICON_GEAR Lidarr ver: v${LIDARR_VERSION_MAJOR:-3} expected"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Checks ━━━
|
||||
# ==============================================================================================
|
||||
check_api "$LIDARR_URL" "Lidarr" 10 || {
|
||||
notify "Lidarr duplicate cleanup aborted on $(hostname) — API unreachable" \
|
||||
"Lidarr Duplicate Cleanup" "warning"
|
||||
exit 1
|
||||
}
|
||||
check_arr_version "$LIDARR_URL" "$LIDARR_API_KEY" "v1" "${LIDARR_VERSION_MAJOR:-3}" "Lidarr" || exit 1
|
||||
|
||||
TMP_DIR=$(mktemp -d)
|
||||
trap 'rm -rf "$TMP_DIR"' EXIT
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Artists (cache-aware — waits out an active rescan rather than trusting a
|
||||
# mid-scan number, falls back to cache if Lidarr's still busy after the strike limit) ━━━
|
||||
# ==============================================================================================
|
||||
ARTISTS=$(arr_get_tracked_data "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")
|
||||
if [[ -z "$ARTISTS" ]]; then
|
||||
warn "Lidarr busy and no usable cache — deferring to next scheduled run"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ARTIST_COUNT=$(echo "$ARTISTS" | jq 'length' 2>/dev/null)
|
||||
if [[ -z "$ARTIST_COUNT" || "$ARTIST_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 artists — aborting to avoid acting on empty data"
|
||||
notify "Lidarr duplicate cleanup aborted on $(hostname) — 0 artists returned" \
|
||||
"Lidarr Duplicate Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
log "$ICON_GEAR Fetched $ARTIST_COUNT artists"
|
||||
|
||||
# Safety Layer — same shared baseline as lidarr_cleanup.sh. A library-wide desync (e.g. mid
|
||||
# full-rescan) can make trackFileCount read far lower than reality for many artists at once —
|
||||
# confirmed 2026-07-16, where this exact scenario would have made the script see both sides
|
||||
# of a genuinely-real duplicate (ROMES) as 0-file phantoms and delete the wrong thing entirely.
|
||||
# Reuses lidarr_cleanup.sh's own baseline file — one shared "is Lidarr's data trustworthy
|
||||
# right now" answer for every script that depends on tracked counts, not a separate opinion
|
||||
# per script.
|
||||
TOTAL_TRACKED=$(echo "$ARTISTS" | jq '[.[].statistics.trackFileCount] | add' 2>/dev/null)
|
||||
check_tracked_count_floor "${TOTAL_TRACKED:-0}" "$LIDARR_TRACKED_COUNT_FILE" "${LIDARR_MIN_TRACKED_PCT:-50}" "Lidarr Duplicate Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Find Case-Insensitive Duplicate Names ━━━
|
||||
# ==============================================================================================
|
||||
DUP_NAMES=$(echo "$ARTISTS" | jq -r '.[].artistName' | tr '[:upper:]' '[:lower:]' | sort | uniq -d)
|
||||
|
||||
if [[ -z "$DUP_NAMES" ]]; then
|
||||
echo "Lidarr — clean ✅ no duplicate artist names"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
DUP_COUNT=$(echo "$DUP_NAMES" | grep -c .)
|
||||
warn "Found $DUP_COUNT duplicate artist name(s)"
|
||||
|
||||
DELETED=0
|
||||
EXCLUDED=0
|
||||
LEFT_ALONE=0
|
||||
FLAGGED=0
|
||||
FLAGGED_NAMES=()
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Resolve Each Duplicate ━━━
|
||||
# ==============================================================================================
|
||||
while IFS= read -r lname; do
|
||||
[[ -z "$lname" ]] && continue
|
||||
|
||||
MEMBERS=$(echo "$ARTISTS" | jq -c --arg n "$lname" '[.[] | select((.artistName|ascii_downcase)==$n)]')
|
||||
|
||||
NONZERO_IDS=()
|
||||
ZERO_MEMBERS_FILE="$TMP_DIR/zero_${RANDOM}.jsonl"
|
||||
: > "$ZERO_MEMBERS_FILE"
|
||||
|
||||
while IFS= read -r member; do
|
||||
[[ -z "$member" ]] && continue
|
||||
fc=$(echo "$member" | jq -r '.statistics.trackFileCount // 0')
|
||||
if [[ "$fc" -gt 0 ]]; then
|
||||
NONZERO_IDS+=("$(echo "$member" | jq -r '.id')")
|
||||
else
|
||||
echo "$member" >> "$ZERO_MEMBERS_FILE"
|
||||
fi
|
||||
done < <(echo "$MEMBERS" | jq -c '.[]')
|
||||
|
||||
if [[ "${#NONZERO_IDS[@]}" -le 1 ]]; then
|
||||
# Simple phantom case — delete every zero-file member, keep the real one (if any)
|
||||
while IFS= read -r zmember; do
|
||||
[[ -z "$zmember" ]] && continue
|
||||
zid=$(echo "$zmember" | jq -r '.id')
|
||||
zname=$(echo "$zmember" | jq -r '.artistName')
|
||||
zmbid=$(echo "$zmember" | jq -r '.foreignArtistId')
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn " DRY RUN — would delete phantom: $zname ($zid) and exclude MBID $zmbid"
|
||||
continue
|
||||
fi
|
||||
|
||||
if curl -sf --max-time 15 -X DELETE \
|
||||
"${LIDARR_URL}/api/v1/artist/${zid}?deleteFiles=false" \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" >/dev/null 2>&1; then
|
||||
(( DELETED++ ))
|
||||
log " $ICON_TRASH Deleted phantom: $zname ($zid)"
|
||||
else
|
||||
warn " Failed to delete phantom: $zname ($zid)"
|
||||
continue
|
||||
fi
|
||||
|
||||
EXCL_PAYLOAD=$(jq -c -n --arg fid "$zmbid" --arg name "$zname" \
|
||||
'{foreignId:$fid, artistName:$name}')
|
||||
if curl -sf --max-time 15 -X POST \
|
||||
"${LIDARR_URL}/api/v1/importlistexclusion" \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" -H "Content-Type: application/json" \
|
||||
-d "$EXCL_PAYLOAD" >/dev/null 2>&1; then
|
||||
(( EXCLUDED++ ))
|
||||
log " Added to import list exclusions: $zmbid"
|
||||
else
|
||||
warn " Could not add exclusion for $zname ($zmbid) — may already exist"
|
||||
fi
|
||||
done < "$ZERO_MEMBERS_FILE"
|
||||
|
||||
# Nudge the surviving real artist so anything stuck on this ambiguity
|
||||
# resolves now rather than waiting for Lidarr's own next check cycle.
|
||||
if [[ "${#NONZERO_IDS[@]}" -eq 1 && "$DRY_RUN" != true ]]; then
|
||||
curl -sf --max-time 15 -X POST \
|
||||
"${LIDARR_URL}/api/v1/command" \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" -H "Content-Type: application/json" \
|
||||
-d "{\"name\":\"RefreshArtist\",\"artistId\":${NONZERO_IDS[0]}}" >/dev/null 2>&1
|
||||
fi
|
||||
else
|
||||
# 2+ members have real content — same name, need to know if it's the same
|
||||
# artist actually split (album overlap) or genuinely different acts (no overlap).
|
||||
#
|
||||
# Titles must be deduped WITHIN each artist before checking overlap ACROSS artists —
|
||||
# a single artist can legitimately list the same album title twice (a reissue, a
|
||||
# deluxe edition under an unchanged title). Confirmed 2026-07-16: Alice Cooper's own
|
||||
# catalog has "Lace and Whiskey" and "School's Out" each listed twice under one
|
||||
# artist ID — treating that as "overlap" false-flagged a genuinely disjoint pair
|
||||
# (band-era vs. solo-era) as ambiguous when it wasn't.
|
||||
OVERLAP=false
|
||||
declare -A GLOBAL_TITLES
|
||||
for nid in "${NONZERO_IDS[@]}"; do
|
||||
declare -A THIS_ARTIST_TITLES
|
||||
while IFS= read -r title; do
|
||||
[[ -z "$title" ]] && continue
|
||||
THIS_ARTIST_TITLES["$title"]=1
|
||||
done < <(curl -sf --max-time 15 "${LIDARR_URL}/api/v1/album?artistId=${nid}" \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" 2>/dev/null | jq -r '.[].title')
|
||||
|
||||
for title in "${!THIS_ARTIST_TITLES[@]}"; do
|
||||
[[ -n "${GLOBAL_TITLES[$title]:-}" ]] && OVERLAP=true
|
||||
GLOBAL_TITLES["$title"]=1
|
||||
done
|
||||
unset THIS_ARTIST_TITLES
|
||||
done
|
||||
unset GLOBAL_TITLES
|
||||
|
||||
if [[ "$OVERLAP" == true ]]; then
|
||||
warn " $ICON_WARN Ambiguous duplicate needs manual review: $lname (both have files, albums overlap)"
|
||||
FLAGGED_NAMES+=("$lname")
|
||||
(( FLAGGED++ ))
|
||||
else
|
||||
log " $lname — different real artists sharing a name, no album overlap, no action"
|
||||
(( LEFT_ALONE++ ))
|
||||
fi
|
||||
fi
|
||||
done <<< "$DUP_NAMES"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY LIDARR DUPLICATE ARTIST CLEANUP SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TRASH Deleted: $DELETED phantom artist(s)"
|
||||
echo "$ICON_GEAR Excluded: $EXCLUDED (blocked from future re-add)"
|
||||
echo "$ICON_SKIP Left alone: $LEFT_ALONE (different real artists, no overlap)"
|
||||
echo "$ICON_WARN Flagged: $FLAGGED (needs manual review)"
|
||||
for n in "${FLAGGED_NAMES[@]}"; do
|
||||
echo " - $n"
|
||||
done
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ "$FLAGGED" -gt 0 ]]; then
|
||||
notify "Lidarr duplicate cleanup on $(hostname) — $DELETED phantom(s) removed, $FLAGGED artist(s) need manual review: ${FLAGGED_NAMES[*]}" \
|
||||
"Lidarr Duplicate Cleanup" "warning"
|
||||
elif [[ "$DELETED" -gt 0 ]]; then
|
||||
echo "$ICON_DONE Status: cleaned $DELETED phantom artist(s), nothing needs review"
|
||||
else
|
||||
echo "$ICON_DONE Status: no action needed"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
exit 0
|
||||
Executable
+668
@@ -0,0 +1,668 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Lidarr Missing Art =========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Fetch missing album and artist artwork for the Lidarr music library. Downloads
|
||||
# only what is absent — never overwrites existing files. Idempotent re-runs are
|
||||
# safe.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Artwork targets per album folder: cover.jpg cdart.png back.jpg
|
||||
# Artwork targets per artist folder: folder.jpg fanart.jpg clearlogo.png banner.jpg
|
||||
#
|
||||
# Sources (tried in order, first success wins):
|
||||
# Album covers: fanart.tv → iTunes fallback
|
||||
# Artist art: fanart.tv → Deezer fallback → Last.fm fallback
|
||||
#
|
||||
# Reads from Lidarr API only — no writes back to Lidarr. Never modifies audio
|
||||
# tags or renames media files. Only writes missing artwork files to existing
|
||||
# album/artist directories.
|
||||
#
|
||||
# Cache-first, both layers (2026-07-17). The artist list comes from the shared tracked-data
|
||||
# cache via arr_get_tracked_data() — fresh (kept warm every 30min by arr_cache_prefill.sh),
|
||||
# live fetch as fallback. The per-artist track-file walk used to build the album→directory
|
||||
# map now reads lidarr_cleanup.sh's write-through cache first (arr_get_cached_items() —
|
||||
# lidarr_cleanup.sh runs earlier in the same nightly window and already does this exact
|
||||
# walk for its own cleanup decisions), falling back to its own live per-artist walk only if
|
||||
# that cache is missing or from outside the current window.
|
||||
#
|
||||
# Negative art cache (2026-07-28). fanart.tv has no cdart/back art for most of the long tail —
|
||||
# roughly 80% of this library — so the nightly run was re-asking about the same ~11.8K albums
|
||||
# every night and getting the same nothing back. Measured 2h41m for 3 images fetched. A flat TSV
|
||||
# in DATA_DIR now remembers "upstream has no <art type> for <mbid>" and skips the API call
|
||||
# entirely until LIDARR_ART_RECHECK_DAYS has passed, so new fanart.tv contributions are still
|
||||
# picked up, just monthly instead of nightly. Only genuine no-art-upstream results are cached —
|
||||
# a failed download of a URL that did exist stays retryable on the next run.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Additive Only
|
||||
# The script only adds missing files — it never overwrites existing artwork
|
||||
# or touches audio files. Re-running after a partial fetch completes exactly
|
||||
# where it left off with no side effects.
|
||||
#
|
||||
# Source Fallback Chain
|
||||
# Multiple sources are tried in order of quality preference. fanart.tv is
|
||||
# primary; fallbacks exist so partial coverage is better than none. A failed
|
||||
# primary never blocks the fallback from running.
|
||||
#
|
||||
# External API Courtesy
|
||||
# Rate limiting and parallel job caps prevent hammering fanart.tv and other
|
||||
# external APIs. Burst behaviour during large initial runs would risk
|
||||
# temporary blocks that break future scheduled fetches.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# acquire_lock — prevents concurrent runs during large library scans
|
||||
# curl + jq check — fail fast if tools missing
|
||||
# API reachability — verified before processing begins
|
||||
# detect_hosts() — exits cleanly if LIDARR_URL empty (HOST2, no Lidarr)
|
||||
# Skip existing — never overwrites, idempotent re-runs are safe
|
||||
# Min file size check — rejects corrupt/placeholder downloads (LIDARR_ART_MIN_SIZE)
|
||||
# Parallel job cap — LIDARR_ART_MAX_PARALLEL — avoids hammering external APIs
|
||||
# Download retries — LIDARR_ART_RETRIES attempts per image before giving up
|
||||
# Rate limiting — LIDARR_ART_SLEEP_BETWEEN between fanart.tv API calls
|
||||
# Negative cache — skips entities whose every missing target is a known upstream miss
|
||||
# Cache expiry on merge — entries past the recheck window are dropped, file stays bounded
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY
|
||||
# Aliased by detect_hosts() — script uses LIDARR_URL / LIDARR_API_KEY
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# FANART_API_KEY — fanart.tv API key
|
||||
# LASTFM_API_KEY — last.fm API key
|
||||
# LIDARR_ART_MIN_SIZE — minimum valid download size in bytes
|
||||
# LIDARR_ART_MAX_PARALLEL — concurrent background download jobs
|
||||
# LIDARR_ART_RETRIES — download retry attempts per image
|
||||
# LIDARR_ART_SLEEP_BETWEEN — seconds between fanart.tv API calls
|
||||
# LIDARR_ART_RECHECK_DAYS — days before re-querying art upstream didn't have
|
||||
# LIDARR_ART_MISS_CACHE — path to the negative cache TSV
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# lidarr_missing_art.sh — fetch all missing artwork
|
||||
# lidarr_missing_art.sh --dry-run — preview without downloading
|
||||
# lidarr_missing_art.sh --log — verbose per-item output
|
||||
# lidarr_missing_art.sh --refresh — ignore the negative cache, re-query everything
|
||||
# lidarr_missing_art.sh --status — show config and exit
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# --refresh ignores the negative cache for this run — use after fanart.tv has had time to gain
|
||||
# new contributions, or to re-prove a miss set by hand. Does not clear the cache; misses found
|
||||
# this run simply overwrite their old stamps.
|
||||
ART_REFRESH=false
|
||||
for _arg in "${PARSED_ARGS[@]}"; do
|
||||
[[ "$_arg" == "--refresh" ]] && ART_REFRESH=true
|
||||
done
|
||||
unset _arg
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for API calls"
|
||||
exit 1
|
||||
fi
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# HOST guard — Lidarr runs on HOST1 only
|
||||
if [[ -z "$LIDARR_URL" ]]; then
|
||||
echo "Lidarr not configured for $MY_ID — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Build path map from MY_ID's Lidarr path map
|
||||
declare -A ARR_PATH_MAP
|
||||
local_path_map_var="${MY_ID}_LIDARR_PATH_MAP"
|
||||
eval "for _key in \"\${!${local_path_map_var}[@]}\"; do
|
||||
ARR_PATH_MAP[\"\$_key\"]=\"\${${local_path_map_var}[\$_key]}\"
|
||||
done"
|
||||
unset _key
|
||||
|
||||
info "$MY_ID ($LOCAL_SERVER_NAME) — tools OK"
|
||||
log "$ICON_GEAR Config: url=${LIDARR_URL}"
|
||||
|
||||
# Defaulted here rather than relying solely on master.conf: Configurations/master.conf is
|
||||
# gitignored, so these reach a node through conf_upgrade seeding them from
|
||||
# Deployment/master.conf.template — which lands in the same window as this script, not before
|
||||
# it. Same guard style as arr_corruption_scan.sh's state file.
|
||||
LIDARR_ART_RECHECK_DAYS="${LIDARR_ART_RECHECK_DAYS:-30}"
|
||||
LIDARR_ART_MISS_CACHE="${LIDARR_ART_MISS_CACHE:-$DATA_DIR/lidarr_art_miss_cache.tsv}"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Lidarr URL: $LIDARR_URL"
|
||||
echo "$ICON_NET Fanart key: $([[ -n "${FANART_API_KEY:-}" ]] && echo "set" || echo "not set")"
|
||||
echo "$ICON_NET LastFM key: $([[ -n "${LASTFM_API_KEY:-}" ]] && echo "set" || echo "not set")"
|
||||
echo "$ICON_GEAR Min size: ${LIDARR_ART_MIN_SIZE} bytes"
|
||||
echo "$ICON_GEAR Parallel: $LIDARR_ART_MAX_PARALLEL jobs"
|
||||
echo "$ICON_RETRY Retries: $LIDARR_ART_RETRIES"
|
||||
echo "$ICON_TIME API sleep: ${LIDARR_ART_SLEEP_BETWEEN}s"
|
||||
echo "$ICON_GEAR Miss cache: $LIDARR_ART_MISS_CACHE ($([[ -f "$LIDARR_ART_MISS_CACHE" ]] && wc -l < "$LIDARR_ART_MISS_CACHE" || echo 0) entries)"
|
||||
echo "$ICON_TIME Recheck: every ${LIDARR_ART_RECHECK_DAYS}d"
|
||||
echo "$ICON_NOTIFY Notify: unRAID=${NOTIFY_UNRAID:-false} Discord=$([[ -n "${MY_DISCORD_WEBHOOK:-}" ]] && echo enabled || echo disabled)"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be written"
|
||||
|
||||
# ── Temp dir for subshell fetch/fail counters ─────────────────────────────────────────────────
|
||||
LIDARR_TMP=$(mktemp -d)
|
||||
trap '_release_all_locks; rm -rf "$LIDARR_TMP"' EXIT
|
||||
touch "$LIDARR_TMP/album_fetches" "$LIDARR_TMP/album_nourl" "$LIDARR_TMP/album_dlfail" \
|
||||
"$LIDARR_TMP/artist_fetches" "$LIDARR_TMP/artist_nourl" "$LIDARR_TMP/artist_dlfail" \
|
||||
"$LIDARR_TMP/album_misses" "$LIDARR_TMP/artist_misses"
|
||||
|
||||
# ── Negative art cache state ──────────────────────────────────────────────────────────────────
|
||||
NOW=$(date +%s)
|
||||
ART_RECHECK_SECS=$(( LIDARR_ART_RECHECK_DAYS * 86400 ))
|
||||
mkdir -p "$(dirname "$LIDARR_ART_MISS_CACHE")"
|
||||
touch "$LIDARR_ART_MISS_CACHE"
|
||||
|
||||
declare -A ART_MISS
|
||||
while IFS=$'\t' read -r _m_key _m_stamp; do
|
||||
[[ -n "$_m_key" ]] && ART_MISS["$_m_key"]="$_m_stamp"
|
||||
done < "$LIDARR_ART_MISS_CACHE"
|
||||
unset _m_key _m_stamp
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FUNCTIONS ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
curl_json() {
|
||||
curl -s --connect-timeout 5 --max-time 20 "$1"
|
||||
}
|
||||
|
||||
job_count() {
|
||||
jobs -rp | wc -l
|
||||
}
|
||||
|
||||
wait_for_slot() {
|
||||
while (( $(job_count) >= LIDARR_ART_MAX_PARALLEL )); do
|
||||
sleep 0.2
|
||||
done
|
||||
}
|
||||
|
||||
# Downloads URL to dest only if dest doesn't exist and downloaded size >= MIN_SIZE.
|
||||
# Returns 0 on success or skip (file already exists), 1 when the source had no URL to offer,
|
||||
# 2 when a URL existed but every download attempt failed.
|
||||
#
|
||||
# The 1-vs-2 split is what makes the negative cache safe: 1 means upstream genuinely has no
|
||||
# such artwork and is worth remembering, 2 means a transient network/CDN problem that must
|
||||
# stay retryable. Caching a 2 would suppress a legitimate retry for LIDARR_ART_RECHECK_DAYS.
|
||||
download_if_valid() {
|
||||
local url="$1"
|
||||
local dest="$2"
|
||||
|
||||
[[ -z "$url" || "$url" == "null" ]] && return 1
|
||||
[[ -f "$dest" ]] && return 0
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
log "DRY RUN — would fetch: $(basename "$dest")"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local tmp="${dest}.tmp"
|
||||
local i
|
||||
for (( i=0; i<=LIDARR_ART_RETRIES; i++ )); do
|
||||
curl -s --connect-timeout 5 --max-time 20 -L -o "$tmp" "$url"
|
||||
local size
|
||||
size=$(stat -c%s "$tmp" 2>/dev/null || echo 0)
|
||||
if (( size > LIDARR_ART_MIN_SIZE )); then
|
||||
mv "$tmp" "$dest"
|
||||
log " Fetched: $(basename "$dest")"
|
||||
return 0
|
||||
fi
|
||||
rm -f "$tmp"
|
||||
sleep 1
|
||||
done
|
||||
|
||||
warn "Failed to fetch: $(basename "$dest")"
|
||||
return 2
|
||||
}
|
||||
|
||||
# ── Negative art cache ────────────────────────────────────────────────────────────────────────
|
||||
# Same shape as arr_corruption_scan.sh's clean-file skip cache: a flat TSV of "<key>\t<epoch>"
|
||||
# in DATA_DIR, slurped into an assoc array once at startup.
|
||||
#
|
||||
# Keyed per (entity, artwork filename) so an album that got cover.jpg from iTunes but has no
|
||||
# cdart upstream still caches the cdart miss alone. MBID is the key where present because it
|
||||
# survives a Lidarr DB rebuild; albums with no MBID fall back to the Lidarr id.
|
||||
art_key() {
|
||||
local mbid="$1" fallback_id="$2"
|
||||
if [[ -n "$mbid" && "$mbid" != "null" ]]; then echo "$mbid"; else echo "lidarrid:$fallback_id"; fi
|
||||
}
|
||||
|
||||
art_miss_fresh() {
|
||||
local stamp="${ART_MISS[$1]:-}"
|
||||
[[ -z "$stamp" ]] && return 1
|
||||
(( NOW - stamp < ART_RECHECK_SECS ))
|
||||
}
|
||||
|
||||
# True when every artwork target still missing from disk is a known-fresh upstream miss —
|
||||
# i.e. this entity cannot possibly gain anything from another round of API calls right now.
|
||||
# This is the check that skips the fanart.tv request entirely, which is where the time goes.
|
||||
art_all_cached() {
|
||||
local dir="$1"; local key="$2"; shift 2
|
||||
local target
|
||||
[[ "$ART_REFRESH" == true ]] && return 1
|
||||
for target in "$@"; do
|
||||
[[ -f "$dir/$target" ]] && continue
|
||||
art_miss_fresh "${key}:${target}" || return 1
|
||||
done
|
||||
return 0
|
||||
}
|
||||
|
||||
# Merges this run's fresh misses into the persistent cache, newest wins per key (fresh file is
|
||||
# read first so awk's first-seen is always the newer stamp). Entries past the recheck window are
|
||||
# dropped rather than carried: they would be re-queried on the next run anyway, so expiring them
|
||||
# here is free and keeps the file from growing without bound as albums leave the library.
|
||||
merge_art_misses() {
|
||||
local fresh="$1" tmp
|
||||
[[ "$DRY_RUN" == true ]] && return 0
|
||||
[[ -s "$fresh" ]] || return 0
|
||||
tmp=$(mktemp)
|
||||
cat "$fresh" "$LIDARR_ART_MISS_CACHE" 2>/dev/null |
|
||||
awk -F'\t' -v cutoff="$(( NOW - ART_RECHECK_SECS ))" \
|
||||
'NF==2 && !seen[$1]++ && $2 >= cutoff' | sort > "$tmp"
|
||||
mv "$tmp" "$LIDARR_ART_MISS_CACHE"
|
||||
}
|
||||
|
||||
# Runs one single-source artwork target and classifies the outcome. Must be called from inside
|
||||
# a fetch subshell — it updates that subshell's _fetches/_nourl/_dlfail/_miss_keys directly.
|
||||
try_single() {
|
||||
local dest="$1" url="$2" key="$3" rc
|
||||
[[ -f "$dest" ]] && return 0
|
||||
download_if_valid "$url" "$dest"; rc=$?
|
||||
case "$rc" in
|
||||
0) (( _fetches++ )) ;;
|
||||
2) (( _dlfail++ )) ;;
|
||||
*) (( _nourl++ )); _miss_keys+="${key}"$'\t'"${NOW}"$'\n' ;;
|
||||
esac
|
||||
}
|
||||
|
||||
deezer_artist_image() {
|
||||
local artist="$1"
|
||||
local query
|
||||
query=$(printf "%s" "$artist" | sed 's/ /+/g')
|
||||
curl_json "https://api.deezer.com/search/artist?q=$query" |
|
||||
jq -r '.data[0].picture_xl // empty' 2>/dev/null
|
||||
}
|
||||
|
||||
lastfm_artist_image() {
|
||||
local artist="$1"
|
||||
local encoded
|
||||
encoded=$(printf "%s" "$artist" | sed 's/ /%20/g')
|
||||
curl_json "https://ws.audioscrobbler.com/2.0/?method=artist.getinfo&artist=$encoded&api_key=$LASTFM_API_KEY&format=json" |
|
||||
jq -r '.artist.image[-1]["#text"] // empty' 2>/dev/null
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Verify Lidarr reachable ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_NET Lidarr API ━━━"
|
||||
|
||||
if ! curl_json "$LIDARR_URL/api/v1/system/status?apikey=$LIDARR_API_KEY" | jq -e '.version' >/dev/null 2>&1; then
|
||||
error "Lidarr API unreachable at $LIDARR_URL — aborting"
|
||||
notify "lidarr_missing_art failed — Lidarr API unreachable on $(hostname)" "Lidarr Missing Art" "warning"
|
||||
exit 1
|
||||
fi
|
||||
info "Lidarr reachable — $LIDARR_URL"
|
||||
|
||||
START=$(date +%s)
|
||||
|
||||
ALBUMS_CHECKED=0
|
||||
ALBUMS_COMPLETE=0
|
||||
ALBUMS_CACHED=0
|
||||
|
||||
ARTISTS_CHECKED=0
|
||||
ARTISTS_COMPLETE=0
|
||||
ARTISTS_CACHED=0
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Build Album Directory Map ━━━
|
||||
# ==============================================================================================
|
||||
# Lidarr's album API never populates .path — derive album dirs from track file paths instead.
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Building Album Directory Map ━━━"
|
||||
|
||||
declare -A ALBUM_DIR_MAP
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (now kept
|
||||
# current every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to
|
||||
# a live fetch when it's stale, and waits out an active rescan before either.
|
||||
_artist_list=$(arr_get_tracked_data "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")
|
||||
_map_artist_count=$(echo "$_artist_list" | jq '. | length')
|
||||
info "Fetching track files for $_map_artist_count artists..."
|
||||
|
||||
# Cache-first for the per-track data too — lidarr_cleanup.sh (runs earlier in the same nightly
|
||||
# window) already does this exact per-artist walk for its own cleanup decisions and writes the
|
||||
# result through via arr_item_cache_write(). Read that instead of repeating the walk; fall
|
||||
# back to the live per-artist walk below only if it's missing or from outside this window.
|
||||
# See arr_item_cache_write()/arr_get_cached_items() in common.sh. (2026-07-17)
|
||||
_cached_tracks=$(arr_get_cached_items "lidarr")
|
||||
if [[ -n "$_cached_tracks" ]]; then
|
||||
info "Using cached track data from lidarr_cleanup.sh — skipping live per-artist walk"
|
||||
while IFS=$'\t' read -r _album_id _track_path; do
|
||||
[[ -z "$_album_id" || -z "$_track_path" || "$_track_path" == "null" ]] && continue
|
||||
# Parameter expansion instead of external dirname — this loop runs once per track
|
||||
# (127K+ on this library), and dirname forks a subprocess per call. Measured
|
||||
# 2026-07-17: ~185x faster (0.39s vs 72.2s for 20K calls) for the identical result.
|
||||
ALBUM_DIR_MAP["$_album_id"]="${_track_path%/*}"
|
||||
done < <(echo "$_cached_tracks" | jq -r '.[] | [(.albumId | tostring), .path] | @tsv' 2>/dev/null)
|
||||
else
|
||||
while IFS= read -r _artist_id; do
|
||||
[[ -z "$_artist_id" ]] && continue
|
||||
while IFS=$'\t' read -r _album_id _track_path; do
|
||||
[[ -z "$_album_id" || -z "$_track_path" || "$_track_path" == "null" ]] && continue
|
||||
ALBUM_DIR_MAP["$_album_id"]="${_track_path%/*}"
|
||||
done < <(curl_json "$LIDARR_URL/api/v1/trackFile?artistId=${_artist_id}&apikey=$LIDARR_API_KEY" | \
|
||||
jq -r '.[] | [(.albumId | tostring), .path] | @tsv' 2>/dev/null)
|
||||
done < <(echo "$_artist_list" | jq -r '.[].id')
|
||||
fi
|
||||
unset _artist_list _map_artist_count _artist_id _album_id _track_path _cached_tracks
|
||||
|
||||
info "Mapped ${#ALBUM_DIR_MAP[@]} albums with local tracks"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Albums ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_EMBY Albums ━━━"
|
||||
|
||||
albums=$(curl_json "$LIDARR_URL/api/v1/album?apikey=$LIDARR_API_KEY")
|
||||
|
||||
if [[ -z "$albums" || "$albums" == "null" ]]; then
|
||||
error "Lidarr album API returned empty — aborting"
|
||||
notify "lidarr_missing_art failed — album API empty on $(hostname)" "Lidarr Missing Art" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
total_albums=$(echo "$albums" | jq '. | length')
|
||||
info "Processing $total_albums albums..."
|
||||
|
||||
while IFS=$'\t' read -r mbid artist_name album_name album_id; do
|
||||
raw_dir="${ALBUM_DIR_MAP[$album_id]:-}"
|
||||
[[ -z "$raw_dir" ]] && continue # not downloaded, skip
|
||||
local_path=$(translate_path "$raw_dir")
|
||||
(( ALBUMS_CHECKED++ ))
|
||||
|
||||
[[ ! -d "$local_path" ]] && continue
|
||||
|
||||
log "[$ALBUMS_CHECKED/$total_albums] $artist_name — $album_name"
|
||||
|
||||
if [[ -f "$local_path/cover.jpg" &&
|
||||
-f "$local_path/cdart.png" &&
|
||||
-f "$local_path/back.jpg" ]]; then
|
||||
(( ALBUMS_COMPLETE++ ))
|
||||
log " complete — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
akey=$(art_key "$mbid" "$album_id")
|
||||
|
||||
if art_all_cached "$local_path" "$akey" cover.jpg cdart.png back.jpg; then
|
||||
(( ALBUMS_CACHED++ ))
|
||||
log " every missing target is a known upstream miss — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
wait_for_slot
|
||||
|
||||
(
|
||||
_fetches=0 _nourl=0 _dlfail=0 _miss_keys=""
|
||||
|
||||
JSON=""
|
||||
if [[ -n "$mbid" && "$mbid" != "null" ]]; then
|
||||
JSON=$(curl_json "http://webservice.fanart.tv/v3/music/albums/$mbid?api_key=$FANART_API_KEY")
|
||||
sleep "$LIDARR_ART_SLEEP_BETWEEN"
|
||||
fi
|
||||
|
||||
# cover.jpg has a two-source chain, so it tracks whether *any* source offered a URL:
|
||||
# only a clean no-URL-anywhere result is cacheable.
|
||||
if [[ ! -f "$local_path/cover.jpg" ]]; then
|
||||
_saw_url=false
|
||||
IMG=$(echo "$JSON" | jq -r '.albums[].albumcover[0].url // empty' 2>/dev/null)
|
||||
download_if_valid "$IMG" "$local_path/cover.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
query=$(printf "%s %s" "$artist_name" "$album_name" | sed 's/ /+/g')
|
||||
itunes=$(curl_json "https://itunes.apple.com/search?term=$query&entity=album&limit=1" |
|
||||
jq -r '.results[0].artworkUrl100 // empty' 2>/dev/null | sed 's/100x100/600x600/')
|
||||
download_if_valid "$itunes" "$local_path/cover.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
elif [[ "$_saw_url" == true ]]; then
|
||||
(( _dlfail++ ))
|
||||
else
|
||||
(( _nourl++ )); _miss_keys+="${akey}:cover.jpg"$'\t'"${NOW}"$'\n'
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
try_single "$local_path/cdart.png" \
|
||||
"$(echo "$JSON" | jq -r '.albums[].cdart[0].url // empty' 2>/dev/null)" \
|
||||
"${akey}:cdart.png"
|
||||
|
||||
try_single "$local_path/back.jpg" \
|
||||
"$(echo "$JSON" | jq -r '.albums[].albumback[0].url // empty' 2>/dev/null)" \
|
||||
"${akey}:back.jpg"
|
||||
|
||||
(( _fetches > 0 )) && printf '1\n' >> "$LIDARR_TMP/album_fetches"
|
||||
(( _nourl > 0 )) && printf '1\n' >> "$LIDARR_TMP/album_nourl"
|
||||
(( _dlfail > 0 )) && printf '1\n' >> "$LIDARR_TMP/album_dlfail"
|
||||
[[ -n "$_miss_keys" ]] && printf '%s' "$_miss_keys" >> "$LIDARR_TMP/album_misses"
|
||||
) &
|
||||
|
||||
done < <(echo "$albums" | jq -r '.[] | [(.foreignAlbumId // ""), (.artist.artistName // ""), (.title // ""), (.id | tostring)] | @tsv')
|
||||
|
||||
wait
|
||||
|
||||
merge_art_misses "$LIDARR_TMP/album_misses"
|
||||
|
||||
ALBUM_FETCHED=$(wc -l < "$LIDARR_TMP/album_fetches" 2>/dev/null || echo 0)
|
||||
ALBUM_NOART=$(wc -l < "$LIDARR_TMP/album_nourl" 2>/dev/null || echo 0)
|
||||
ALBUM_DLFAIL=$(wc -l < "$LIDARR_TMP/album_dlfail" 2>/dev/null || echo 0)
|
||||
ALBUM_MISSING=$(( ALBUMS_CHECKED - ALBUMS_COMPLETE ))
|
||||
info "Checked: $ALBUMS_CHECKED | Complete: $ALBUMS_COMPLETE | Needed art: $ALBUM_MISSING | Cached-skip: $ALBUMS_CACHED | Fetched: $ALBUM_FETCHED | No art upstream: $ALBUM_NOART | Fetch failed: $ALBUM_DLFAIL"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Artists ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_EMBY Artists ━━━"
|
||||
|
||||
# Cache-first — see the earlier album-map fetch above for why (arr_get_tracked_data serves
|
||||
# the shared cache when fresh, live fetch as fallback, waits out an active rescan first).
|
||||
artists=$(arr_get_tracked_data "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1")
|
||||
|
||||
if [[ -z "$artists" || "$artists" == "null" ]]; then
|
||||
error "Lidarr artist API returned empty — aborting"
|
||||
notify "lidarr_missing_art failed — artist API empty on $(hostname)" "Lidarr Missing Art" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
total_artists=$(echo "$artists" | jq '. | length')
|
||||
info "Processing $total_artists artists..."
|
||||
|
||||
while IFS=$'\t' read -r local_path mbid name; do
|
||||
local_path=$(translate_path "$local_path")
|
||||
(( ARTISTS_CHECKED++ ))
|
||||
|
||||
[[ ! -d "$local_path" ]] && continue
|
||||
[[ -z "$mbid" || "$mbid" == "null" ]] && continue
|
||||
|
||||
log "[$ARTISTS_CHECKED/$total_artists] $name"
|
||||
|
||||
if [[ -f "$local_path/folder.jpg" &&
|
||||
-f "$local_path/fanart.jpg" &&
|
||||
-f "$local_path/clearlogo.png" &&
|
||||
-f "$local_path/banner.jpg" ]]; then
|
||||
(( ARTISTS_COMPLETE++ ))
|
||||
log " complete — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
if art_all_cached "$local_path" "$mbid" folder.jpg fanart.jpg clearlogo.png banner.jpg; then
|
||||
(( ARTISTS_CACHED++ ))
|
||||
log " every missing target is a known upstream miss — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
wait_for_slot
|
||||
|
||||
(
|
||||
_fetches=0 _nourl=0 _dlfail=0 _miss_keys=""
|
||||
|
||||
JSON=$(curl_json "http://webservice.fanart.tv/v3/music/$mbid?api_key=$FANART_API_KEY")
|
||||
sleep "$LIDARR_ART_SLEEP_BETWEEN"
|
||||
|
||||
# folder.jpg walks fanart → Deezer → Last.fm; only a no-URL result from all three is
|
||||
# cacheable, same reasoning as the album cover chain.
|
||||
if [[ ! -f "$local_path/folder.jpg" ]]; then
|
||||
_saw_url=false
|
||||
IMG=$(echo "$JSON" | jq -r '.artistthumb[0].url // empty' 2>/dev/null)
|
||||
download_if_valid "$IMG" "$local_path/folder.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
IMG=$(deezer_artist_image "$name")
|
||||
download_if_valid "$IMG" "$local_path/folder.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
IMG=$(lastfm_artist_image "$name")
|
||||
download_if_valid "$IMG" "$local_path/folder.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
elif [[ "$_saw_url" == true ]]; then
|
||||
(( _dlfail++ ))
|
||||
else
|
||||
(( _nourl++ )); _miss_keys+="${mbid}:folder.jpg"$'\t'"${NOW}"$'\n'
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/fanart.jpg" ]]; then
|
||||
_saw_url=false
|
||||
IMG=$(echo "$JSON" | jq -r '.artistbackground[0].url // empty' 2>/dev/null)
|
||||
download_if_valid "$IMG" "$local_path/fanart.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
IMG=$(deezer_artist_image "$name")
|
||||
download_if_valid "$IMG" "$local_path/fanart.jpg"; _rc=$?
|
||||
(( _rc == 2 )) && _saw_url=true
|
||||
if (( _rc == 0 )); then
|
||||
(( _fetches++ ))
|
||||
elif [[ "$_saw_url" == true ]]; then
|
||||
(( _dlfail++ ))
|
||||
else
|
||||
(( _nourl++ )); _miss_keys+="${mbid}:fanart.jpg"$'\t'"${NOW}"$'\n'
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
try_single "$local_path/clearlogo.png" \
|
||||
"$(echo "$JSON" | jq -r '.hdmusiclogo[0].url // empty' 2>/dev/null)" \
|
||||
"${mbid}:clearlogo.png"
|
||||
|
||||
try_single "$local_path/banner.jpg" \
|
||||
"$(echo "$JSON" | jq -r '.musicbanner[0].url // empty' 2>/dev/null)" \
|
||||
"${mbid}:banner.jpg"
|
||||
|
||||
(( _fetches > 0 )) && printf '1\n' >> "$LIDARR_TMP/artist_fetches"
|
||||
(( _nourl > 0 )) && printf '1\n' >> "$LIDARR_TMP/artist_nourl"
|
||||
(( _dlfail > 0 )) && printf '1\n' >> "$LIDARR_TMP/artist_dlfail"
|
||||
[[ -n "$_miss_keys" ]] && printf '%s' "$_miss_keys" >> "$LIDARR_TMP/artist_misses"
|
||||
) &
|
||||
|
||||
done < <(echo "$artists" | jq -r '.[] | [.path, .foreignArtistId, .artistName] | @tsv')
|
||||
|
||||
wait
|
||||
|
||||
merge_art_misses "$LIDARR_TMP/artist_misses"
|
||||
|
||||
ARTIST_FETCHED=$(wc -l < "$LIDARR_TMP/artist_fetches" 2>/dev/null || echo 0)
|
||||
ARTIST_NOART=$(wc -l < "$LIDARR_TMP/artist_nourl" 2>/dev/null || echo 0)
|
||||
ARTIST_DLFAIL=$(wc -l < "$LIDARR_TMP/artist_dlfail" 2>/dev/null || echo 0)
|
||||
ARTIST_MISSING=$(( ARTISTS_CHECKED - ARTISTS_COMPLETE ))
|
||||
info "Checked: $ARTISTS_CHECKED | Complete: $ARTISTS_COMPLETE | Needed art: $ARTIST_MISSING | Cached-skip: $ARTISTS_CACHED | Fetched: $ARTIST_FETCHED | No art upstream: $ARTIST_NOART | Fetch failed: $ARTIST_DLFAIL"
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY LIDARR MISSING ART SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $((END - START)))"
|
||||
echo "$ICON_EMBY Albums: $ALBUMS_CHECKED checked | $ALBUMS_COMPLETE complete | $ALBUM_FETCHED fetched"
|
||||
echo "$ICON_SKIP Albums: $ALBUMS_CACHED skipped (cached miss) | $ALBUM_NOART no art upstream | $ALBUM_DLFAIL fetch failed"
|
||||
echo "$ICON_EMBY Artists: $ARTISTS_CHECKED checked | $ARTISTS_COMPLETE complete | $ARTIST_FETCHED fetched"
|
||||
echo "$ICON_SKIP Artists: $ARTISTS_CACHED skipped (cached miss) | $ARTIST_NOART no art upstream | $ARTIST_DLFAIL fetch failed"
|
||||
echo "$ICON_GEAR Cache: $([[ -f "$LIDARR_ART_MISS_CACHE" ]] && wc -l < "$LIDARR_ART_MISS_CACHE" || echo 0) known upstream misses | recheck every ${LIDARR_ART_RECHECK_DAYS}d"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo "$ICON_WARN Status: DRY RUN — no files written"
|
||||
else
|
||||
echo "$ICON_DONE Status: $ICON_SUCCESS ALL DONE"
|
||||
notify "Lidarr art fetch complete on $(hostname) — ${ALBUMS_CHECKED} albums, ${ARTISTS_CHECKED} artists processed" "Lidarr Missing Art" "normal"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
exit 0
|
||||
Executable
+493
@@ -0,0 +1,493 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================= Lidarr Release Fixer ===========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Fix Lidarr albums where the wrong MusicBrainz release edition was selected,
|
||||
# causing on-disk files to appear as unimported despite being present.
|
||||
#
|
||||
# Root cause: Lidarr tracks one specific release edition per album using
|
||||
# foreignReleaseId. When this doesn't match the MUSICBRAINZ_ALBUMID embedded
|
||||
# in the actual files, track ID lookup fails and RescanFolders imports 0 tracks
|
||||
# even with perfect, fully tagged files.
|
||||
#
|
||||
# Fix: Read the MusicBrainz Album ID from the first FLAC or MP3 found in each
|
||||
# album directory, look that release up in Lidarr's known releases for that
|
||||
# album, switch monitored=true to the correct one, then queue a RefreshArtist
|
||||
# to re-sync track IDs and trigger Lidarr's own import scan.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# For each monitored album with 0 tracked files:
|
||||
# 1. Locate the album directory under the artist's root path (title glob match)
|
||||
# 2. Find the first FLAC or MP3 file in that directory
|
||||
# 3. Read MUSICBRAINZ_ALBUMID from the file's tags
|
||||
# 4. Fetch the album's available releases from Lidarr
|
||||
# 5. If the file's release exists and differs from Lidarr's current selection:
|
||||
# — PUT the album with the correct release set to monitored=true
|
||||
# — Mark the artist for a RefreshArtist command
|
||||
#
|
||||
# RefreshArtist is batched — one per artist, even if multiple albums were fixed.
|
||||
# Lidarr handles the post-refresh rescan and import automatically.
|
||||
#
|
||||
# Cache-first artist list (2026-07-17) — comes from the shared tracked-data cache via
|
||||
# arr_get_tracked_data(), fresh (kept warm every 30min by arr_cache_prefill.sh), live fetch
|
||||
# as fallback. Everything below (per-album release lookups) still fetches live — that data
|
||||
# isn't part of what's cached.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Fixer Before Cleanup
|
||||
# Runs before lidarr_cleanup.sh in the daily job list. The strike system in
|
||||
# cleanup provides a protection window, but correcting releases first means
|
||||
# files that were wrongly treated as orphans get imported rather than aged out.
|
||||
#
|
||||
# No Blind Fixes
|
||||
# Only switches to a release that is already in Lidarr's known release list for
|
||||
# that album. If the file's MBID isn't a recognized release, the album is skipped
|
||||
# rather than guessed at. False corrections are worse than leaving it alone.
|
||||
#
|
||||
# Files Are Never Touched
|
||||
# This script only modifies the Lidarr database record (release selection). Audio
|
||||
# files, tags, and directory structure are read-only. All actual importing is
|
||||
# handled by Lidarr after RefreshArtist runs.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# check_api — API reachability verified before any processing
|
||||
# check_arr_version — aborts if Lidarr major version doesn't match LIDARR_VERSION_MAJOR
|
||||
# artist count > 0 — aborts if artist fetch returns empty (API anomaly guard)
|
||||
# release list validation — only fixes to releases explicitly known to Lidarr for that album
|
||||
# acquire_lock "skip" — skips if another instance is already running
|
||||
# detect_hosts() — exits cleanly on hosts without Lidarr configured
|
||||
# curl + jq + perl check — fail fast if any required tool is missing
|
||||
#
|
||||
# ==============================================================================================
|
||||
# TAG READING
|
||||
# ==============================================================================================
|
||||
#
|
||||
# FLAC — Vorbis comment block (block type 4), key MUSICBRAINZ_ALBUMID
|
||||
# MP3 — ID3v2 TXXX frame, description "MusicBrainz Album Id"
|
||||
# Supports Latin-1, UTF-8 (enc 0/3) and UTF-16 (enc 1/2) encodings
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY / HOST1_LIDARR_MUSIC_ROOT
|
||||
# HOST1_LIDARR_PATH_MAP — container path → host path translation
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# LIDARR_RELEASE_FIXER_ENABLED — set false to disable without removing from job list
|
||||
# LIDARR_VERSION_MAJOR — expected Lidarr major version for API safety check
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# lidarr_release_fixer.sh — normal run
|
||||
# lidarr_release_fixer.sh --dry-run — preview, no API writes
|
||||
# lidarr_release_fixer.sh --log — verbose output
|
||||
# lidarr_release_fixer.sh --status — show config and exit
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
parse_args "$@"
|
||||
|
||||
# ── Setup ─────────────────────────────────────────────────────────────────────────────────────
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
for _cmd in curl jq perl; do
|
||||
if ! command -v "$_cmd" >/dev/null 2>&1; then
|
||||
error "$_cmd not found — required"
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
unset _cmd
|
||||
|
||||
acquire_lock "skip"
|
||||
trap "_release_all_locks" EXIT
|
||||
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "${LIDARR_URL:-}" ]] || [[ -z "${LIDARR_API_KEY:-}" ]]; then
|
||||
info "Lidarr not configured on $MY_ID — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ "${LIDARR_RELEASE_FIXER_ENABLED:-true}" == "false" ]]; then
|
||||
info "LIDARR_RELEASE_FIXER_ENABLED=false — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
declare -A ARR_PATH_MAP
|
||||
local_path_map_var="${MY_ID}_LIDARR_PATH_MAP"
|
||||
eval "for key in \"\${!${local_path_map_var}[@]}\"; do
|
||||
ARR_PATH_MAP[\"\$key\"]=\"\${${local_path_map_var}[\$key]}\"
|
||||
done"
|
||||
|
||||
require_var LIDARR_URL
|
||||
require_var LIDARR_API_KEY
|
||||
require_var LIDARR_MUSIC_ROOT
|
||||
|
||||
# ── Status ────────────────────────────────────────────────────────────────────────────────────
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Lidarr URL: $LIDARR_URL"
|
||||
echo "$ICON_GEAR Music root: $LIDARR_MUSIC_ROOT"
|
||||
echo "$ICON_GEAR Enabled: ${LIDARR_RELEASE_FIXER_ENABLED:-true}"
|
||||
echo "$ICON_GEAR Dry run: $DRY_RUN"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ── API helpers ───────────────────────────────────────────────────────────────────────────────
|
||||
lidarr_api() {
|
||||
local endpoint="$1"
|
||||
local response http_code body
|
||||
response=$(curl -sf --max-time 30 \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${LIDARR_URL}/api/v1/${endpoint}" 2>/dev/null)
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
if [[ "$http_code" != "200" ]]; then
|
||||
error "Lidarr API HTTP $http_code for: $endpoint"
|
||||
return 1
|
||||
fi
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
lidarr_api_put() {
|
||||
local endpoint="$1"
|
||||
local payload="$2"
|
||||
local http_code
|
||||
http_code=$(curl -sf --max-time 30 -X PUT \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$payload" \
|
||||
-w "%{http_code}" -o /dev/null \
|
||||
"${LIDARR_URL}/api/v1/${endpoint}" 2>/dev/null)
|
||||
if [[ "$http_code" != "202" ]] && [[ "$http_code" != "200" ]]; then
|
||||
error "Lidarr PUT HTTP $http_code for: $endpoint"
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
lidarr_api_post() {
|
||||
local endpoint="$1"
|
||||
local payload="$2"
|
||||
curl -sf --max-time 30 -X POST \
|
||||
-H "X-Api-Key: $LIDARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$payload" \
|
||||
-o /dev/null \
|
||||
"${LIDARR_URL}/api/v1/${endpoint}" 2>/dev/null
|
||||
}
|
||||
|
||||
# ── Tag readers ───────────────────────────────────────────────────────────────────────────────
|
||||
# Read MUSICBRAINZ_ALBUMID from a FLAC file's Vorbis comment block (block type 4)
|
||||
read_flac_mbid() {
|
||||
perl -e '
|
||||
open(my $fh, "<:raw", $ARGV[0]) or exit;
|
||||
read($fh, my $magic, 4); substr($magic, 0, 4) eq "fLaC" or exit;
|
||||
while (1) {
|
||||
read($fh, my $hdr, 4) == 4 or last;
|
||||
my $w = unpack("N", $hdr);
|
||||
my $last = ($w >> 31) & 1;
|
||||
my $type = ($w >> 24) & 0x7f;
|
||||
my $len = $w & 0xffffff;
|
||||
read($fh, my $data, $len);
|
||||
if ($type == 4) {
|
||||
my $pos = 0;
|
||||
my $vl = unpack("V", substr($data, $pos, 4)); $pos += 4 + $vl;
|
||||
my $n = unpack("V", substr($data, $pos, 4)); $pos += 4;
|
||||
for (1..$n) {
|
||||
my $cl = unpack("V", substr($data, $pos, 4)); $pos += 4;
|
||||
my $c = substr($data, $pos, $cl); $pos += $cl;
|
||||
if ($c =~ /^MUSICBRAINZ_ALBUMID=(.+)$/i) { print "$1\n"; exit; }
|
||||
}
|
||||
}
|
||||
last if $last;
|
||||
}
|
||||
' "$1" 2>/dev/null
|
||||
}
|
||||
|
||||
# Read MusicBrainz Album Id from an MP3's ID3v2 TXXX frame.
|
||||
# Handles Latin-1/UTF-8 (enc 0/3) with single-null separator and
|
||||
# UTF-16 (enc 1/2) by stripping null bytes and pattern-matching.
|
||||
read_mp3_mbid() {
|
||||
perl -e '
|
||||
open(my $fh, "<:raw", $ARGV[0]) or exit;
|
||||
read($fh, my $hdr, 10) == 10 or exit;
|
||||
substr($hdr, 0, 3) eq "ID3" or exit;
|
||||
my $ver = ord(substr($hdr, 3, 1));
|
||||
my $sz = 0; $sz = ($sz << 7) | ord($_) for split //, substr($hdr, 6, 4);
|
||||
read($fh, my $data, $sz) == $sz or exit;
|
||||
my $pos = 0;
|
||||
while ($pos + 10 <= $sz) {
|
||||
my $id = substr($data, $pos, 4); $pos += 4;
|
||||
last unless $id =~ /^[A-Z][A-Z0-9]{3}$/;
|
||||
my $fs;
|
||||
if ($ver >= 4) {
|
||||
my $n = 0; $n = ($n << 7) | ord($_) for split //, substr($data, $pos, 4);
|
||||
$fs = $n;
|
||||
} else {
|
||||
$fs = unpack("N", substr($data, $pos, 4));
|
||||
}
|
||||
$pos += 6;
|
||||
last if $fs < 1 || $pos + $fs > $sz;
|
||||
if ($id eq "TXXX") {
|
||||
my $enc = ord(substr($data, $pos, 1));
|
||||
my $body = substr($data, $pos + 1, $fs - 1);
|
||||
if ($enc == 1 || $enc == 2) {
|
||||
# UTF-16: strip BOM, collapse to ASCII, pattern match
|
||||
$body =~ s/^\xff\xfe|^\xfe\xff//;
|
||||
(my $flat = $body) =~ s/\x00//g;
|
||||
if ($flat =~ /^MusicBrainz Album Id([0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12})/i) {
|
||||
print "$1\n"; exit;
|
||||
}
|
||||
} else {
|
||||
my ($desc, $val) = split /\x00/, $body, 2;
|
||||
if (defined $desc && lc($desc) eq "musicbrainz album id" &&
|
||||
defined $val &&
|
||||
$val =~ /^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$/i) {
|
||||
print "$val\n"; exit;
|
||||
}
|
||||
}
|
||||
}
|
||||
$pos += $fs;
|
||||
}
|
||||
' "$1" 2>/dev/null
|
||||
}
|
||||
|
||||
read_file_mbid() {
|
||||
local file="$1"
|
||||
case "${file##*.}" in
|
||||
[Ff][Ll][Aa][Cc]) read_flac_mbid "$file" ;;
|
||||
[Mm][Pp]3) read_mp3_mbid "$file" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# ── Safety checks ─────────────────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
|
||||
if ! check_api "$LIDARR_URL" "Lidarr" 10; then
|
||||
notify "Lidarr release fixer aborted on $(hostname) — API unreachable" \
|
||||
"Lidarr Release Fixer" "warning"
|
||||
exit 1
|
||||
fi
|
||||
check_arr_version "$LIDARR_URL" "$LIDARR_API_KEY" "v1" "$LIDARR_VERSION_MAJOR" "Lidarr" || exit 1
|
||||
info "API reachable and version OK"
|
||||
|
||||
# ── Fetch all artists and build path cache ────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching artist paths ━━━"
|
||||
|
||||
declare -A ARTIST_PATH_CACHE
|
||||
declare -A ARTIST_NAME_CACHE
|
||||
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (now kept
|
||||
# current every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to
|
||||
# a live fetch when it's stale, and waits out an active rescan before either.
|
||||
ALL_ARTISTS=$(arr_get_tracked_data "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1") || {
|
||||
error "Failed to fetch artists"
|
||||
exit 1
|
||||
}
|
||||
|
||||
while IFS= read -r artist; do
|
||||
aid=$(echo "$artist" | jq -r '.id')
|
||||
apath=$(echo "$artist" | jq -r '.path // empty')
|
||||
aname=$(echo "$artist" | jq -r '.artistName // empty')
|
||||
[[ -n "$apath" ]] && ARTIST_PATH_CACHE[$aid]=$(translate_path "$apath")
|
||||
[[ -n "$aname" ]] && ARTIST_NAME_CACHE[$aid]="$aname"
|
||||
done < <(echo "$ALL_ARTISTS" | jq -c '.[]' 2>/dev/null)
|
||||
unset ALL_ARTISTS
|
||||
|
||||
ARTIST_COUNT="${#ARTIST_PATH_CACHE[@]}"
|
||||
info "Loaded $ARTIST_COUNT artists"
|
||||
|
||||
if [[ "$ARTIST_COUNT" -eq 0 ]]; then
|
||||
error "No artists returned — aborting"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ── Fetch zero-file albums ────────────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching zero-file monitored albums ━━━"
|
||||
|
||||
ALL_ALBUMS=$(lidarr_api "album") || {
|
||||
error "Failed to fetch albums"
|
||||
exit 1
|
||||
}
|
||||
|
||||
ZERO_FILE_ALBUMS=$(echo "$ALL_ALBUMS" | jq -c \
|
||||
'[.[] | select(.monitored == true and .statistics.trackFileCount == 0)]' 2>/dev/null)
|
||||
unset ALL_ALBUMS
|
||||
|
||||
TOTAL_ZERO=$(echo "$ZERO_FILE_ALBUMS" | jq 'length' 2>/dev/null)
|
||||
info "Monitored albums with 0 tracked files: $TOTAL_ZERO"
|
||||
|
||||
# ── Process albums ────────────────────────────────────────────────────────────────────────────
|
||||
START=$(date +%s)
|
||||
FIXED=0
|
||||
SKIPPED_NO_DIR=0
|
||||
SKIPPED_NO_FILE=0
|
||||
SKIPPED_NO_MBID=0
|
||||
SKIPPED_NO_MATCH=0
|
||||
SKIPPED_CORRECT=0
|
||||
ERRORS=0
|
||||
declare -A ARTISTS_TO_REFRESH
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR Processing albums ━━━"
|
||||
|
||||
while IFS= read -r album; do
|
||||
album_id=$(echo "$album" | jq -r '.id')
|
||||
album_title=$(echo "$album" | jq -r '.title')
|
||||
artist_id=$(echo "$album" | jq -r '.artistId')
|
||||
artist_name="${ARTIST_NAME_CACHE[$artist_id]:-artist $artist_id}"
|
||||
|
||||
log "Checking: $artist_name — $album_title (album $album_id)"
|
||||
|
||||
artist_host_path="${ARTIST_PATH_CACHE[$artist_id]:-}"
|
||||
if [[ -z "$artist_host_path" ]] || [[ ! -d "$artist_host_path" ]]; then
|
||||
log "Artist dir not found: $artist_host_path"
|
||||
(( SKIPPED_NO_DIR++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Find album directory — case-insensitive prefix match on title
|
||||
album_dir=$(find "$artist_host_path" -maxdepth 1 -type d -iname "${album_title}*" \
|
||||
2>/dev/null | head -1)
|
||||
if [[ -z "$album_dir" ]]; then
|
||||
log "Album dir not found: $album_title"
|
||||
(( SKIPPED_NO_DIR++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Find first FLAC or MP3
|
||||
music_file=$(find "$album_dir" -maxdepth 2 -type f \
|
||||
\( -iname "*.flac" -o -iname "*.mp3" \) 2>/dev/null | head -1)
|
||||
if [[ -z "$music_file" ]]; then
|
||||
log "No FLAC or MP3 in: $album_dir"
|
||||
(( SKIPPED_NO_FILE++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Read MBID from file tags
|
||||
file_mbid=$(read_file_mbid "$music_file")
|
||||
if [[ -z "$file_mbid" ]]; then
|
||||
log "No MBID tag in: $music_file"
|
||||
(( SKIPPED_NO_MBID++ ))
|
||||
continue
|
||||
fi
|
||||
log "File MBID: $file_mbid"
|
||||
|
||||
# Fetch full album JSON (includes releases array)
|
||||
album_full=$(lidarr_api "album/${album_id}") || { (( ERRORS++ )); sleep 0.2; continue; }
|
||||
sleep 0.1
|
||||
|
||||
# Currently selected release
|
||||
current_release_id=$(echo "$album_full" | jq -r \
|
||||
'.releases[] | select(.monitored == true) | .foreignReleaseId' 2>/dev/null | head -1)
|
||||
|
||||
if [[ "$current_release_id" == "$file_mbid" ]]; then
|
||||
log "Release already correct: $file_mbid"
|
||||
(( SKIPPED_CORRECT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Verify file's MBID is a known release for this album
|
||||
release_known=$(echo "$album_full" | jq -r \
|
||||
--arg rid "$file_mbid" \
|
||||
'.releases[] | select(.foreignReleaseId == $rid) | .foreignReleaseId' 2>/dev/null)
|
||||
if [[ -z "$release_known" ]]; then
|
||||
log "File MBID $file_mbid not in Lidarr release list for: $album_title"
|
||||
(( SKIPPED_NO_MATCH++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Build updated album object with correct release selected
|
||||
updated_album=$(echo "$album_full" | jq \
|
||||
--arg target "$file_mbid" \
|
||||
'.releases = [.releases[] | .monitored = (.foreignReleaseId == $target)]' 2>/dev/null)
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would fix: $artist_name — $album_title"
|
||||
warn " $current_release_id → $file_mbid"
|
||||
(( FIXED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
if lidarr_api_put "album/${album_id}" "$updated_album"; then
|
||||
warn "$ICON_GEAR Fixed: $artist_name — $album_title"
|
||||
warn " $current_release_id → $file_mbid"
|
||||
ARTISTS_TO_REFRESH[$artist_id]="$artist_id"
|
||||
(( FIXED++ ))
|
||||
else
|
||||
(( ERRORS++ ))
|
||||
fi
|
||||
sleep 0.2
|
||||
|
||||
done < <(echo "$ZERO_FILE_ALBUMS" | jq -c '.[]')
|
||||
|
||||
# ── Queue RefreshArtist for all fixed artists ─────────────────────────────────────────────────
|
||||
if [[ "${#ARTISTS_TO_REFRESH[@]}" -gt 0 ]] && [[ "$DRY_RUN" == false ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Queuing RefreshArtist ━━━"
|
||||
for artist_id in "${!ARTISTS_TO_REFRESH[@]}"; do
|
||||
if lidarr_api_post "command" \
|
||||
"{\"name\": \"RefreshArtist\", \"artistId\": ${artist_id}}"; then
|
||||
log "Queued RefreshArtist for: ${ARTIST_NAME_CACHE[$artist_id]:-$artist_id}"
|
||||
else
|
||||
error "Failed to queue RefreshArtist for artist $artist_id"
|
||||
fi
|
||||
sleep 0.1
|
||||
done
|
||||
info "Refresh queued for ${#ARTISTS_TO_REFRESH[@]} artist(s)"
|
||||
fi
|
||||
|
||||
# ── Summary ───────────────────────────────────────────────────────────────────────────────────
|
||||
END=$(date +%s)
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY LIDARR RELEASE FIXER SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Candidates: $TOTAL_ZERO zero-file albums"
|
||||
echo "$ICON_DONE Fixed: $FIXED"
|
||||
echo "$ICON_SKIP Already correct: $SKIPPED_CORRECT"
|
||||
echo "$ICON_SKIP No dir on disk: $SKIPPED_NO_DIR"
|
||||
echo "$ICON_SKIP No music file: $SKIPPED_NO_FILE"
|
||||
echo "$ICON_SKIP No MBID tag: $SKIPPED_NO_MBID"
|
||||
echo "$ICON_SKIP MBID not in list: $SKIPPED_NO_MATCH"
|
||||
[[ "$ERRORS" -gt 0 ]] && echo " Errors: $ERRORS"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes made"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
if [[ "$FIXED" -gt 0 ]] && [[ "$DRY_RUN" == false ]]; then
|
||||
notify "Lidarr release fixer on $(hostname) — corrected $FIXED album release(s)" \
|
||||
"Lidarr Release Fixer" "normal"
|
||||
fi
|
||||
|
||||
exit 0
|
||||
+89
-40
@@ -16,7 +16,7 @@
|
||||
# Goal: 0–5 meaningful Lidarr adds per week, not bulk imports.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# FLOW
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Fetch play completions from Emby activity log (last LOOKBACK_DAYS days)
|
||||
@@ -33,6 +33,11 @@
|
||||
# 8. Score candidates: affinity + breadth + popularity + quality
|
||||
# 9. Take top MAX_ADDS above threshold → add to Lidarr
|
||||
#
|
||||
# Step 7's "already in Lidarr" check reads the shared tracked-data cache via
|
||||
# arr_get_tracked_data() (cache-first, live fallback, 2026-07-17) instead of a live fetch —
|
||||
# this runs weekly right after arr_full_rescan.sh, so it's reading the genuine post-rescan
|
||||
# snapshot arr_full_rescan.sh just wrote.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# SCORING MODEL
|
||||
# ==============================================================================================
|
||||
@@ -53,34 +58,45 @@
|
||||
# Max adds: LIDARR_DISCOVERY_MAX_ADDS (default 5) caps each stage
|
||||
#
|
||||
# ==============================================================================================
|
||||
# REQUIREMENTS
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Last.fm API key — required for both stages
|
||||
# Configure HOST*_LASTFM_API_KEY in host*.conf
|
||||
# host*.conf (aliased by detect_hosts())
|
||||
#
|
||||
# HOST*_LASTFM_API_KEY Required for both stages — without it the run exits cleanly
|
||||
# rather than adding anything unscored.
|
||||
# LIDARR_URL / LIDARR_API_KEY Target arr
|
||||
# EMBY_URL / EMBY_API_KEY Play history source for seed selection
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# LIDARR_DISCOVERY_THRESHOLD Score required to accept a candidate (0-100)
|
||||
# LIDARR_DISCOVERY_LOOKBACK_DAYS Emby play history window
|
||||
# LIDARR_DISCOVERY_MIN_PLAYS Min plays in the window before an artist is evaluated
|
||||
# LIDARR_DISCOVERY_USER_CAP_PCT Max % of the play score any one user can contribute,
|
||||
# so a single heavy listener cannot drive the library
|
||||
# LIDARR_DISCOVERY_MAX_ADDS Hard cap on artists added per run
|
||||
# LIDARR_DISCOVERY_REJECT_COOLDOWN Days before a rejected artist is re-evaluated
|
||||
# LIDARR_DISCOVERY_HISTORY Decision history DB — accepted and rejected
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION (master.conf)
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# LIDARR_DISCOVERY_THRESHOLD — minimum score for Stage 1 seeds and Stage 2 adds (default: 70)
|
||||
# LIDARR_DISCOVERY_LOOKBACK_DAYS — Emby play history window in days (default: 7)
|
||||
# LIDARR_DISCOVERY_MIN_PLAYS — min plays to be evaluated in Stage 1 (default: 3)
|
||||
# LIDARR_DISCOVERY_MAX_ADDS — max seeds (Stage 1) and max adds (Stage 2) (default: 5)
|
||||
# LIDARR_DISCOVERY_USER_CAP_PCT — max % any one user contributes to play weight (default: 35)
|
||||
# LIDARR_DISCOVERY_REJECT_COOLDOWN — days before re-evaluating a Stage 2 reject (default: 30)
|
||||
# LIDARR_DISCOVERY_HISTORY — history/state file path
|
||||
# Playback as Intent Signal
|
||||
# What users actually listen to is a stronger signal than what they follow or
|
||||
# own. The scoring model weights demonstrated listening behaviour — recency,
|
||||
# play count, user breadth — over passive library membership.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
# Selective by Design
|
||||
# 0–5 adds per week is the target, not bulk imports. A high score threshold
|
||||
# combined with MAX_ADDS ensures only high-confidence recommendations are
|
||||
# acted on. Volume is not the goal — meaningful discovery is.
|
||||
#
|
||||
# playback_aware_lidarr_discovery.sh — normal run
|
||||
# playback_aware_lidarr_discovery.sh --dry-run — score and rank, no Lidarr changes
|
||||
# playback_aware_lidarr_discovery.sh --log — verbose output
|
||||
# playback_aware_lidarr_discovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: weekly (WEEKLY_MAINTENANCE_SCRIPTS in master.conf)
|
||||
# Two-Stage Filtering
|
||||
# Stage 1 rejects weak seeds before they drive Stage 2. Low-quality seeds
|
||||
# produce low-quality similar-artist recommendations. Filtering at the seed
|
||||
# stage improves the entire output, not just the top of the list.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
@@ -110,6 +126,29 @@
|
||||
# runs. Safe to delete — next run starts fresh with no memory.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION (master.conf)
|
||||
# ==============================================================================================
|
||||
#
|
||||
# LIDARR_DISCOVERY_THRESHOLD — minimum score for Stage 1 seeds and Stage 2 adds (default: 70)
|
||||
# LIDARR_DISCOVERY_LOOKBACK_DAYS — Emby play history window in days (default: 7)
|
||||
# LIDARR_DISCOVERY_MIN_PLAYS — min plays to be evaluated in Stage 1 (default: 3)
|
||||
# LIDARR_DISCOVERY_MAX_ADDS — max seeds (Stage 1) and max adds (Stage 2) (default: 5)
|
||||
# LIDARR_DISCOVERY_USER_CAP_PCT — max % any one user contributes to play weight (default: 35)
|
||||
# LIDARR_DISCOVERY_REJECT_COOLDOWN — days before re-evaluating a Stage 2 reject (default: 30)
|
||||
# LIDARR_DISCOVERY_HISTORY — history/state file path
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# playback_aware_lidarr_discovery.sh — normal run
|
||||
# playback_aware_lidarr_discovery.sh --dry-run — score and rank, no Lidarr changes
|
||||
# playback_aware_lidarr_discovery.sh --log — verbose output
|
||||
# playback_aware_lidarr_discovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: weekly (WEEKLY_MAINTENANCE_SCRIPTS in master.conf)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
@@ -158,6 +197,8 @@ USER_CAP_PCT="${LIDARR_DISCOVERY_USER_CAP_PCT:-35}"
|
||||
REJECT_COOLDOWN="${LIDARR_DISCOVERY_REJECT_COOLDOWN:-30}"
|
||||
HISTORY_FILE="${LIDARR_DISCOVERY_HISTORY:-${DATA_DIR}/lidarr_discovery_history.db}"
|
||||
|
||||
log "$ICON_GEAR Config: threshold=${THRESHOLD} lookback=${LOOKBACK_DAYS}d min-plays=${MIN_PLAYS} max-adds=${MAX_ADDS} user-cap=${USER_CAP_PCT}% reject-cooldown=${REJECT_COOLDOWN}d"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no artists will be added to Lidarr"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -186,18 +227,6 @@ fi
|
||||
# ── API HELPERS ───────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
_emby_api() {
|
||||
local endpoint="$1"
|
||||
local response http_code body
|
||||
response=$(curl -sf --max-time 15 \
|
||||
-H "X-Emby-Token: $EMBY_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${EMBY_URL}/${endpoint}" 2>/dev/null)
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
[[ "$http_code" != "200" ]] && { error "Emby API HTTP $http_code: $endpoint"; return 1; }
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
_lidarr_get() {
|
||||
curl -sf --max-time 20 \
|
||||
@@ -329,7 +358,7 @@ CUTOFF_ISO=$(date -d "${LOOKBACK_DAYS} days ago" '+%Y-%m-%dT%H:%M:%SZ')
|
||||
TODAY_EPOCH=$(date +%s)
|
||||
TODAY=$(date +%Y-%m-%d)
|
||||
|
||||
ACTIVITY_JSON=$(_emby_api "System/ActivityLog/Entries?MinDate=${CUTOFF_ISO}&Limit=5000") || {
|
||||
ACTIVITY_JSON=$(emby_api "System/ActivityLog/Entries?MinDate=${CUTOFF_ISO}&Limit=5000") || {
|
||||
error "Could not fetch Emby activity log"
|
||||
exit 1
|
||||
}
|
||||
@@ -507,21 +536,41 @@ fi
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Existing Libraries ━━━"
|
||||
|
||||
LIDARR_ARTISTS_JSON=$(_lidarr_get "artist") || { error "Could not fetch Lidarr artists"; exit 1; }
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (kept current
|
||||
# every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to a live
|
||||
# fetch when it's stale, and waits out an active rescan before either.
|
||||
LIDARR_ARTISTS_JSON=$(arr_get_tracked_data "lidarr" "$LIDARR_URL" "$LIDARR_API_KEY" "v1") || { error "Could not fetch Lidarr artists"; exit 1; }
|
||||
LIDARR_NAMES=$(echo "$LIDARR_ARTISTS_JSON" | jq -r '.[].artistName' 2>/dev/null)
|
||||
LIDARR_COUNT=$(echo "$LIDARR_NAMES" | grep -c . 2>/dev/null || echo 0)
|
||||
LIDARR_COUNT=$(echo "$LIDARR_NAMES" | grep -c . 2>/dev/null || true)
|
||||
log "$LIDARR_COUNT artists in Lidarr"
|
||||
|
||||
EMBY_LIBRARY_JSON=$(_emby_api "Items?IncludeItemTypes=MusicAlbum&Recursive=true&Fields=AlbumArtists&Limit=10000") || {
|
||||
EMBY_LIBRARY_JSON=$(emby_api "Items?IncludeItemTypes=MusicAlbum&Recursive=true&Fields=AlbumArtists&Limit=10000") || {
|
||||
warn "Could not fetch Emby artist library — skipping Emby filter"
|
||||
EMBY_ARTIST_NAMES=""
|
||||
}
|
||||
EMBY_ARTIST_NAMES=$(echo "$EMBY_LIBRARY_JSON" | jq -r '.Items[] | .AlbumArtists[]?.Name' 2>/dev/null)
|
||||
EMBY_ARTIST_COUNT=$(echo "$EMBY_ARTIST_NAMES" | grep -c . 2>/dev/null || echo 0)
|
||||
EMBY_ARTIST_COUNT=$(echo "$EMBY_ARTIST_NAMES" | grep -c . 2>/dev/null || true)
|
||||
log "$EMBY_ARTIST_COUNT album artists in Emby library"
|
||||
|
||||
_in_lidarr() { echo "$LIDARR_NAMES" | grep -iq "^${1}$"; }
|
||||
_in_emby_library(){ [[ -n "$EMBY_ARTIST_NAMES" ]] && echo "$EMBY_ARTIST_NAMES" | grep -iq "^${1}$"; }
|
||||
# MusicBrainz's canonical name for some artists (e.g. "blink‐182") uses a Unicode
|
||||
# hyphen/dash rather than plain ASCII "-". Last.fm's candidate names are plain ASCII,
|
||||
# so an exact-string match against Lidarr/Emby's names silently misses these artists
|
||||
# every time — they never register as "already known" and get retried (and rejected
|
||||
# as duplicates) on every future run. Normalize both sides before comparing.
|
||||
_normalize_dashes() {
|
||||
local n="$1"
|
||||
n="${n//‐/-}" # U+2010 HYPHEN
|
||||
n="${n//‑/-}" # U+2011 NON-BREAKING HYPHEN
|
||||
n="${n//‒/-}" # U+2012 FIGURE DASH
|
||||
n="${n//–/-}" # U+2013 EN DASH
|
||||
n="${n//—/-}" # U+2014 EM DASH
|
||||
echo "$n"
|
||||
}
|
||||
LIDARR_NAMES=$(_normalize_dashes "$LIDARR_NAMES")
|
||||
EMBY_ARTIST_NAMES=$(_normalize_dashes "$EMBY_ARTIST_NAMES")
|
||||
|
||||
_in_lidarr() { echo "$LIDARR_NAMES" | grep -iq "^$(_normalize_dashes "$1")$"; }
|
||||
_in_emby_library(){ [[ -n "$EMBY_ARTIST_NAMES" ]] && echo "$EMBY_ARTIST_NAMES" | grep -iq "^$(_normalize_dashes "$1")$"; }
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Score Stage 2 Candidates ━━━
|
||||
+76
-71
@@ -16,7 +16,7 @@
|
||||
# Goal: 0–5 meaningful Radarr adds per run, not bulk imports.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# FLOW
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Fetch recently watched movies from Emby (SEED_LIBRARIES, last LOOKBACK_DAYS days)
|
||||
@@ -33,6 +33,11 @@
|
||||
# 7. Score candidates: breadth + TMDB rating + vote count
|
||||
# 8. Take top MAX_ADDS above threshold → add to Radarr
|
||||
#
|
||||
# Step 6's "already in Radarr" check reads the shared tracked-data cache via
|
||||
# arr_get_tracked_data() (cache-first, live fallback, 2026-07-17) instead of a live fetch —
|
||||
# this runs weekly right after arr_full_rescan.sh, so it's reading the genuine post-rescan
|
||||
# snapshot arr_full_rescan.sh just wrote.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# SCORING MODEL
|
||||
# ==============================================================================================
|
||||
@@ -51,37 +56,47 @@
|
||||
# Max adds: RADARR_DISCOVERY_MAX_ADDS (default 5)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# REQUIREMENTS
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# TMDB API key — required for Stage 2 recommendations
|
||||
# Configure HOST*_TMDB_API_KEY in host*.conf
|
||||
# Free key at: https://www.themoviedb.org/settings/api
|
||||
# host*.conf (aliased by detect_hosts())
|
||||
#
|
||||
# HOST*_TMDB_API_KEY Required for Stage 2 recommendations. Free key at
|
||||
# https://www.themoviedb.org/settings/api — without it the run
|
||||
# exits cleanly rather than adding anything unscored.
|
||||
# RADARR_URL / RADARR_API_KEY Target arr
|
||||
# EMBY_URL / EMBY_API_KEY Play history source for seed selection
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# RADARR_DISCOVERY_THRESHOLD Score required to accept a candidate (0-100)
|
||||
# RADARR_DISCOVERY_LOOKBACK_DAYS Emby watch history window
|
||||
# RADARR_DISCOVERY_MAX_SEEDS Max seed movies taken from Stage 1
|
||||
# RADARR_DISCOVERY_MAX_ADDS Hard cap on movies added per run
|
||||
# RADARR_DISCOVERY_MIN_VOTE_COUNT Min TMDB votes for a candidate to be considered
|
||||
# RADARR_DISCOVERY_MIN_RATING Min TMDB vote_average × 10
|
||||
# RADARR_DISCOVERY_REJECT_COOLDOWN Days before a rejected movie is re-evaluated
|
||||
# RADARR_DISCOVERY_SEED_LIBRARIES Emby libraries to draw seed movies from
|
||||
# RADARR_DISCOVERY_HISTORY Decision history DB — accepted and rejected
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION (master.conf)
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# RADARR_DISCOVERY_THRESHOLD — minimum score to add a candidate (default: 52)
|
||||
# RADARR_DISCOVERY_LOOKBACK_DAYS — Emby watch history window in days (default: 30)
|
||||
# RADARR_DISCOVERY_MAX_SEEDS — max seed movies from Stage 1 (default: 5)
|
||||
# RADARR_DISCOVERY_MAX_ADDS — max movies to add per run (default: 5)
|
||||
# RADARR_DISCOVERY_MIN_VOTE_COUNT — min TMDB votes for a candidate (default: 100)
|
||||
# RADARR_DISCOVERY_MIN_RATING — min TMDB vote_average × 10 (default: 60 = 6.0/10)
|
||||
# RADARR_DISCOVERY_REJECT_COOLDOWN — days before re-evaluating a rejected movie (default: 60)
|
||||
# RADARR_DISCOVERY_SEED_LIBRARIES — Emby library names to draw seeds from (default: ("Movies"))
|
||||
# RADARR_DISCOVERY_HISTORY — history/state file path
|
||||
# Playback as Intent Signal
|
||||
# Recently watched movies are a stronger signal than what is in the library or
|
||||
# on watchlists. The scoring model weights demonstrated viewing behaviour —
|
||||
# recency, rating, vote confidence — over passive ownership.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
# Selective by Design
|
||||
# 0–5 adds per run is the target, not bulk imports. A score threshold combined
|
||||
# with MAX_ADDS ensures only high-confidence recommendations are acted on.
|
||||
# Volume is not the goal — meaningful discovery is.
|
||||
#
|
||||
# playback_aware_radarr_discovery.sh — normal run
|
||||
# playback_aware_radarr_discovery.sh --dry-run — score and rank, no Radarr changes
|
||||
# playback_aware_radarr_discovery.sh --log — verbose output
|
||||
# playback_aware_radarr_discovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: weekly (WEEKLY_MAINTENANCE_SCRIPTS in master.conf)
|
||||
# Two-Stage Filtering
|
||||
# Stage 1 rejects weak seeds before they drive Stage 2. Low-quality or
|
||||
# low-confidence watched movies produce poor recommendations. Filtering at
|
||||
# the seed stage improves the entire output.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
@@ -111,11 +126,35 @@
|
||||
# runs. Safe to delete — next run starts fresh with no memory.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION (master.conf)
|
||||
# ==============================================================================================
|
||||
#
|
||||
# RADARR_DISCOVERY_THRESHOLD — minimum score to add a candidate (default: 52)
|
||||
# RADARR_DISCOVERY_LOOKBACK_DAYS — Emby watch history window in days (default: 30)
|
||||
# RADARR_DISCOVERY_MAX_SEEDS — max seed movies from Stage 1 (default: 5)
|
||||
# RADARR_DISCOVERY_MAX_ADDS — max movies to add per run (default: 5)
|
||||
# RADARR_DISCOVERY_MIN_VOTE_COUNT — min TMDB votes for a candidate (default: 100)
|
||||
# RADARR_DISCOVERY_MIN_RATING — min TMDB vote_average × 10 (default: 60 = 6.0/10)
|
||||
# RADARR_DISCOVERY_REJECT_COOLDOWN — days before re-evaluating a rejected movie (default: 60)
|
||||
# RADARR_DISCOVERY_SEED_LIBRARIES — Emby library names to draw seeds from (default: ("Movies"))
|
||||
# RADARR_DISCOVERY_HISTORY — history/state file path
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# playback_aware_radarr_discovery.sh — normal run
|
||||
# playback_aware_radarr_discovery.sh --dry-run — score and rank, no Radarr changes
|
||||
# playback_aware_radarr_discovery.sh --log — verbose output
|
||||
# playback_aware_radarr_discovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: weekly (WEEKLY_MAINTENANCE_SCRIPTS in master.conf)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
source "$SCRIPT_DIR/../Kernel/decision_engine.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
@@ -162,9 +201,11 @@ REJECT_COOLDOWN="${RADARR_DISCOVERY_REJECT_COOLDOWN:-60}"
|
||||
HISTORY_FILE="${RADARR_DISCOVERY_HISTORY:-${DATA_DIR}/radarr_discovery_history.db}"
|
||||
SEED_LIBRARIES=("${RADARR_DISCOVERY_SEED_LIBRARIES[@]:-Movies}")
|
||||
|
||||
log "$ICON_GEAR Config: threshold=${THRESHOLD} lookback=${LOOKBACK_DAYS}d max-seeds=${MAX_SEEDS} max-adds=${MAX_ADDS} min-votes=${MIN_VOTE_COUNT} min-rating=${MIN_RATING} reject-cooldown=${REJECT_COOLDOWN}d"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no movies will be added to Radarr"
|
||||
|
||||
_fmt_rating() { local v="${1:-0}"; echo "${v::-1}.${v: -1}" 2>/dev/null || echo "$v"; }
|
||||
# _fmt_rating() — provided by common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
@@ -194,17 +235,6 @@ fi
|
||||
# ── API HELPERS ───────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
_emby_api() {
|
||||
local response http_code body
|
||||
response=$(curl -sf --max-time 20 \
|
||||
-H "X-Emby-Token: $EMBY_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${EMBY_URL}/${1}" 2>/dev/null)
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
[[ "$http_code" != "200" ]] && { error "Emby API HTTP $http_code: $1"; return 1; }
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
_radarr_get() {
|
||||
curl -sf --max-time 20 \
|
||||
@@ -265,35 +295,7 @@ _freq_score() {
|
||||
fi
|
||||
}
|
||||
|
||||
# vote_avg_int = vote_average × 10 as integer (e.g. 7.8 → 78)
|
||||
_rating_score_s2() {
|
||||
local v="$1"
|
||||
if (( v >= 80 )); then echo 40
|
||||
elif (( v >= 75 )); then echo 32
|
||||
elif (( v >= 70 )); then echo 25
|
||||
elif (( v >= 65 )); then echo 18
|
||||
elif (( v >= 60 )); then echo 12
|
||||
else echo 5
|
||||
fi
|
||||
}
|
||||
|
||||
_votes_score() {
|
||||
local c="$1"
|
||||
if (( c >= 10000 )); then echo 20
|
||||
elif (( c >= 5000 )); then echo 15
|
||||
elif (( c >= 1000 )); then echo 10
|
||||
elif (( c >= 200 )); then echo 5
|
||||
else echo 2
|
||||
fi
|
||||
}
|
||||
|
||||
_breadth_score() {
|
||||
local seeds="$1"
|
||||
if (( seeds >= 3 )); then echo 40
|
||||
elif (( seeds == 2 )); then echo 25
|
||||
else echo 10
|
||||
fi
|
||||
}
|
||||
# _rating_score_s2(), _votes_score(), _breadth_score() — provided by common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Emby Libraries + Recently Watched Movies ━━━
|
||||
@@ -306,7 +308,7 @@ echo "$ICON_HOST $MY_ID ($LOCAL_SERVER_NAME) | lookback: ${LOOKBACK_DAYS}d | thr
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Emby Watch History ━━━"
|
||||
|
||||
LIBRARIES_JSON=$(_emby_api "Library/VirtualFolders") || { error "Could not fetch Emby libraries"; exit 1; }
|
||||
LIBRARIES_JSON=$(emby_api "Library/VirtualFolders") || { error "Could not fetch Emby libraries"; exit 1; }
|
||||
CUTOFF_ISO=$(date -d "${LOOKBACK_DAYS} days ago" '+%Y-%m-%dT%H:%M:%SZ')
|
||||
TODAY_EPOCH=$(date +%s)
|
||||
TODAY=$(date +%Y-%m-%d)
|
||||
@@ -321,7 +323,7 @@ for lib_name in "${SEED_LIBRARIES[@]}"; do
|
||||
continue
|
||||
fi
|
||||
log "Indexing library: $lib_name (ItemId: $lib_id)"
|
||||
LIB_JSON=$(_emby_api "Items?ParentId=${lib_id}&IncludeItemTypes=Movie&Recursive=true&Fields=ProviderIds&Limit=10000") || {
|
||||
LIB_JSON=$(emby_api "Items?ParentId=${lib_id}&IncludeItemTypes=Movie&Recursive=true&Fields=ProviderIds&Limit=10000") || {
|
||||
warn "Could not index library: $lib_name"
|
||||
continue
|
||||
}
|
||||
@@ -335,7 +337,7 @@ log "${#EMBY_TMDB_IDS[@]} movies in Emby TMDB index"
|
||||
# Each entry includes an ItemId — batch-fetch those items to determine type (Movie vs. Music/TV).
|
||||
# SortBy=DatePlayed on Items requires UserId context and errors without one; the activity log
|
||||
# is server-scoped and doesn't have that limitation.
|
||||
ACTIVITY_JSON=$(_emby_api "System/ActivityLog/Entries?MinDate=${CUTOFF_ISO}&Limit=5000") || {
|
||||
ACTIVITY_JSON=$(emby_api "System/ActivityLog/Entries?MinDate=${CUTOFF_ISO}&Limit=5000") || {
|
||||
error "Could not fetch Emby activity log"
|
||||
exit 1
|
||||
}
|
||||
@@ -366,7 +368,7 @@ if [[ "${#ITEM_PLAYS[@]}" -gt 0 ]]; then
|
||||
for (( _b=0; _b<${#ALL_ITEM_IDS[@]}; _b+=BATCH_SIZE )); do
|
||||
BATCH=("${ALL_ITEM_IDS[@]:_b:BATCH_SIZE}")
|
||||
IDS_CSV=$(printf '%s,' "${BATCH[@]}"); IDS_CSV="${IDS_CSV%,}"
|
||||
ITEMS_DETAIL=$(_emby_api "Items?Ids=${IDS_CSV}&Fields=ProviderIds,Type&Limit=200") || {
|
||||
ITEMS_DETAIL=$(emby_api "Items?Ids=${IDS_CSV}&Fields=ProviderIds,Type&Limit=200") || {
|
||||
warn "Could not fetch item batch starting at $_b"
|
||||
continue
|
||||
}
|
||||
@@ -504,7 +506,10 @@ fi
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Existing Libraries ━━━"
|
||||
|
||||
RADARR_MOVIES_JSON=$(_radarr_get "movie") || { error "Could not fetch Radarr library"; exit 1; }
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (kept current
|
||||
# every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to a live
|
||||
# fetch when it's stale, and waits out an active rescan before either.
|
||||
RADARR_MOVIES_JSON=$(arr_get_tracked_data "radarr" "$RADARR_URL" "$RADARR_API_KEY" "v3") || { error "Could not fetch Radarr library"; exit 1; }
|
||||
|
||||
declare -A RADARR_TMDB # tmdb_id → 1
|
||||
while read -r tmdb_id; do
|
||||
+79
-78
@@ -17,7 +17,7 @@
|
||||
# Goal: 0–3 meaningful Sonarr adds per run, not bulk imports.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# FLOW
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Fetch all Series from Emby SONARR_EMBY_LIBRARIES — build TMDB+TVDB index
|
||||
@@ -37,6 +37,11 @@
|
||||
# 11. Take top MAX_ADDS above threshold
|
||||
# 12. Get TVDB ID via TMDB external_ids → Sonarr lookup → add + trigger SeriesSearch
|
||||
#
|
||||
# Step 9's "already in Sonarr" check reads the shared tracked-data cache via
|
||||
# arr_get_tracked_data() (cache-first, live fallback, 2026-07-17) instead of a live fetch —
|
||||
# this runs weekly right after arr_full_rescan.sh, so it's reading the genuine post-rescan
|
||||
# snapshot arr_full_rescan.sh just wrote.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# SCORING MODEL
|
||||
# ==============================================================================================
|
||||
@@ -58,39 +63,48 @@
|
||||
# Max adds: SONARR_DISCOVERY_MAX_ADDS (default 3) — TV is a larger commitment than movies
|
||||
#
|
||||
# ==============================================================================================
|
||||
# REQUIREMENTS
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# TMDB API key — required for Stage 2 recommendations and external_ids lookup
|
||||
# Configure HOST*_TMDB_API_KEY in host*.conf
|
||||
# Free key at: https://www.themoviedb.org/settings/api
|
||||
# host*.conf (aliased by detect_hosts())
|
||||
#
|
||||
# HOST*_TMDB_API_KEY Required for Stage 2 recommendations and external_ids lookup.
|
||||
# Free key at https://www.themoviedb.org/settings/api — without it
|
||||
# the run exits cleanly rather than adding anything unscored.
|
||||
# SONARR_URL / SONARR_API_KEY Target arr
|
||||
# EMBY_URL / EMBY_API_KEY Play history source for seed selection
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# SONARR_DISCOVERY_THRESHOLD Score required to accept a candidate (0-100)
|
||||
# SONARR_DISCOVERY_LOOKBACK_DAYS Emby episode play history window
|
||||
# SONARR_DISCOVERY_MAX_SEEDS Max seed series taken from Stage 1
|
||||
# SONARR_DISCOVERY_MAX_ADDS Hard cap on shows added per run
|
||||
# SONARR_DISCOVERY_MIN_VOTE_COUNT Min TMDB votes for a candidate to be considered
|
||||
# SONARR_DISCOVERY_MIN_RATING Min TMDB vote_average × 10
|
||||
# SONARR_DISCOVERY_REJECT_COOLDOWN Days before a rejected show is re-evaluated
|
||||
# SONARR_DISCOVERY_USER_EPISODE_CAP Max episodes one user contributes to seed volume
|
||||
# SONARR_DISCOVERY_MONITOR_MODE Sonarr monitor mode on add
|
||||
# SONARR_DISCOVERY_HISTORY Decision history DB — accepted and rejected
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION (master.conf)
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# SONARR_DISCOVERY_THRESHOLD — minimum score to add a candidate (default: 52)
|
||||
# SONARR_DISCOVERY_LOOKBACK_DAYS — Emby watch history window in days (default: 14)
|
||||
# SONARR_DISCOVERY_MAX_SEEDS — max seed series from Stage 1 (default: 5)
|
||||
# SONARR_DISCOVERY_MAX_ADDS — max shows to add per run (default: 3)
|
||||
# SONARR_DISCOVERY_MIN_VOTE_COUNT — min TMDB votes for a candidate (default: 50)
|
||||
# SONARR_DISCOVERY_MIN_RATING — min TMDB vote_average × 10 (default: 65 = 6.5/10)
|
||||
# SONARR_DISCOVERY_REJECT_COOLDOWN — days before re-evaluating a rejected show (default: 60)
|
||||
# SONARR_DISCOVERY_USER_EPISODE_CAP — max episodes per user in seed scoring (default: 8)
|
||||
# SONARR_DISCOVERY_MONITOR_MODE — Sonarr monitor mode on add: "all" or "future" (default: "all")
|
||||
# SONARR_EMBY_LIBRARIES — Emby library names to draw seeds from
|
||||
# SONARR_DISCOVERY_HISTORY — history/state file path
|
||||
# Playback as Intent Signal
|
||||
# Recently watched episodes are a stronger signal than what is in the library.
|
||||
# User diversity across a series is weighted above a single user binge —
|
||||
# broad household interest is a better predictor of a good addition than
|
||||
# one person's session.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
# Selective by Design
|
||||
# 0–3 adds per run is the target. TV is a larger commitment than movies —
|
||||
# a lower MAX_ADDS cap reflects that. Volume is not the goal.
|
||||
#
|
||||
# playback_aware_sonarr_discovery.sh — normal run
|
||||
# playback_aware_sonarr_discovery.sh --dry-run — score and rank, no Sonarr changes
|
||||
# playback_aware_sonarr_discovery.sh --log — verbose output
|
||||
# playback_aware_sonarr_discovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: weekly (WEEKLY_MAINTENANCE_SCRIPTS in master.conf)
|
||||
# Two-Stage Filtering
|
||||
# Stage 1 rejects weak seeds before they drive Stage 2. A poorly-watched
|
||||
# or niche series produces poor recommendations. Filtering at the seed
|
||||
# stage improves the entire output.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
@@ -125,11 +139,37 @@
|
||||
# runs. Safe to delete — next run starts fresh with no memory.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION (master.conf)
|
||||
# ==============================================================================================
|
||||
#
|
||||
# SONARR_DISCOVERY_THRESHOLD — minimum score to add a candidate (default: 52)
|
||||
# SONARR_DISCOVERY_LOOKBACK_DAYS — Emby watch history window in days (default: 14)
|
||||
# SONARR_DISCOVERY_MAX_SEEDS — max seed series from Stage 1 (default: 5)
|
||||
# SONARR_DISCOVERY_MAX_ADDS — max shows to add per run (default: 3)
|
||||
# SONARR_DISCOVERY_MIN_VOTE_COUNT — min TMDB votes for a candidate (default: 50)
|
||||
# SONARR_DISCOVERY_MIN_RATING — min TMDB vote_average × 10 (default: 65 = 6.5/10)
|
||||
# SONARR_DISCOVERY_REJECT_COOLDOWN — days before re-evaluating a rejected show (default: 60)
|
||||
# SONARR_DISCOVERY_USER_EPISODE_CAP — max episodes per user in seed scoring (default: 8)
|
||||
# SONARR_DISCOVERY_MONITOR_MODE — Sonarr monitor mode on add: "all" or "future" (default: "all")
|
||||
# SONARR_EMBY_LIBRARIES — Emby library names to draw seeds from
|
||||
# SONARR_DISCOVERY_HISTORY — history/state file path
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# playback_aware_sonarr_discovery.sh — normal run
|
||||
# playback_aware_sonarr_discovery.sh --dry-run — score and rank, no Sonarr changes
|
||||
# playback_aware_sonarr_discovery.sh --log — verbose output
|
||||
# playback_aware_sonarr_discovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: weekly (WEEKLY_MAINTENANCE_SCRIPTS in master.conf)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
source "$SCRIPT_DIR/../Kernel/decision_engine.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
@@ -182,9 +222,11 @@ USER_EPISODE_CAP="${SONARR_DISCOVERY_USER_EPISODE_CAP:-8}"
|
||||
MONITOR_MODE="${SONARR_DISCOVERY_MONITOR_MODE:-all}"
|
||||
HISTORY_FILE="${SONARR_DISCOVERY_HISTORY:-${DATA_DIR}/sonarr_discovery_history.db}"
|
||||
|
||||
log "$ICON_GEAR Config: threshold=${THRESHOLD} lookback=${LOOKBACK_DAYS}d max-seeds=${MAX_SEEDS} max-adds=${MAX_ADDS} min-votes=${MIN_VOTE_COUNT} min-rating=${MIN_RATING} reject-cooldown=${REJECT_COOLDOWN}d monitor=${MONITOR_MODE}"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no shows will be added to Sonarr"
|
||||
|
||||
_fmt_rating() { local v="${1:-0}"; echo "${v::-1}.${v: -1}" 2>/dev/null || echo "$v"; }
|
||||
# _fmt_rating() — provided by common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
@@ -216,17 +258,6 @@ fi
|
||||
# ── API HELPERS ───────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
_emby_api() {
|
||||
local response http_code body
|
||||
response=$(curl -sf --max-time 30 \
|
||||
-H "X-Emby-Token: $EMBY_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${EMBY_URL}/${1}" 2>/dev/null)
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
[[ "$http_code" != "200" ]] && { error "Emby API HTTP $http_code: $1"; return 1; }
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
_sonarr_get() {
|
||||
curl -sf --max-time 30 \
|
||||
@@ -305,37 +336,7 @@ _volume_score() {
|
||||
fi
|
||||
}
|
||||
|
||||
# Stage 2: TMDB vote_average × 10 (0-40)
|
||||
_rating_score_s2() {
|
||||
local v="$1"
|
||||
if (( v >= 80 )); then echo 40
|
||||
elif (( v >= 75 )); then echo 32
|
||||
elif (( v >= 70 )); then echo 25
|
||||
elif (( v >= 65 )); then echo 18
|
||||
elif (( v >= 60 )); then echo 12
|
||||
else echo 5
|
||||
fi
|
||||
}
|
||||
|
||||
# Stage 2: vote count (0-20)
|
||||
_votes_score() {
|
||||
local c="$1"
|
||||
if (( c >= 10000 )); then echo 20
|
||||
elif (( c >= 5000 )); then echo 15
|
||||
elif (( c >= 1000 )); then echo 10
|
||||
elif (( c >= 200 )); then echo 5
|
||||
else echo 2
|
||||
fi
|
||||
}
|
||||
|
||||
# Stage 2: seed breadth (0-40)
|
||||
_breadth_score() {
|
||||
local seeds="$1"
|
||||
if (( seeds >= 3 )); then echo 40
|
||||
elif (( seeds == 2 )); then echo 25
|
||||
else echo 10
|
||||
fi
|
||||
}
|
||||
# _rating_score_s2(), _votes_score(), _breadth_score() — provided by common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Emby Series Library ━━━
|
||||
@@ -348,7 +349,7 @@ echo "$ICON_HOST $MY_ID ($LOCAL_SERVER_NAME) | lookback: ${LOOKBACK_DAYS}d | thr
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Emby Series Library ━━━"
|
||||
|
||||
LIBRARIES_JSON=$(_emby_api "Library/VirtualFolders") || { error "Could not fetch Emby libraries"; exit 1; }
|
||||
LIBRARIES_JSON=$(emby_api "Library/VirtualFolders") || { error "Could not fetch Emby libraries"; exit 1; }
|
||||
CUTOFF_ISO=$(date -d "${LOOKBACK_DAYS} days ago" '+%Y-%m-%dT%H:%M:%SZ')
|
||||
TODAY_EPOCH=$(date +%s)
|
||||
TODAY=$(date +%Y-%m-%d)
|
||||
@@ -369,7 +370,7 @@ for lib_name in "${SONARR_EMBY_LIBRARIES[@]}"; do
|
||||
continue
|
||||
fi
|
||||
log "Scanning library: $lib_name (ItemId: $lib_id)"
|
||||
LIB_JSON=$(_emby_api "Items?ParentId=${lib_id}&IncludeItemTypes=Series&Recursive=true&Fields=ProviderIds&Limit=5000") || {
|
||||
LIB_JSON=$(emby_api "Items?ParentId=${lib_id}&IncludeItemTypes=Series&Recursive=true&Fields=ProviderIds&Limit=5000") || {
|
||||
warn "Could not fetch series from library: $lib_name"
|
||||
continue
|
||||
}
|
||||
@@ -404,7 +405,7 @@ echo "━━━ $ICON_SYNC Emby Watch History ━━━"
|
||||
declare -A ITEM_USER_PLAYS # "episode_item_id|user_id" → play count
|
||||
declare -A ITEM_LAST_PLAY # episode_item_id → ISO date of most recent play
|
||||
|
||||
ACTIVITY_JSON=$(_emby_api "System/ActivityLog/Entries?MinDate=${CUTOFF_ISO}&Limit=5000") || {
|
||||
ACTIVITY_JSON=$(emby_api "System/ActivityLog/Entries?MinDate=${CUTOFF_ISO}&Limit=5000") || {
|
||||
error "Could not fetch Emby activity log"
|
||||
exit 1
|
||||
}
|
||||
@@ -448,7 +449,7 @@ if [[ "${#ALL_EPISODE_IDS[@]}" -gt 0 ]]; then
|
||||
for (( _b=0; _b<${#ALL_EPISODE_IDS[@]}; _b+=BATCH_SIZE )); do
|
||||
BATCH=("${ALL_EPISODE_IDS[@]:_b:BATCH_SIZE}")
|
||||
IDS_CSV=$(printf '%s,' "${BATCH[@]}"); IDS_CSV="${IDS_CSV%,}"
|
||||
ITEMS_DETAIL=$(_emby_api "Items?Ids=${IDS_CSV}&Fields=SeriesId,Type&Limit=200") || {
|
||||
ITEMS_DETAIL=$(emby_api "Items?Ids=${IDS_CSV}&Fields=SeriesId,Type&Limit=200") || {
|
||||
warn "Could not fetch episode batch starting at $_b"
|
||||
continue
|
||||
}
|
||||
@@ -612,7 +613,10 @@ fi
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Existing Libraries ━━━"
|
||||
|
||||
SONARR_SERIES_JSON=$(_sonarr_get "series") || { error "Could not fetch Sonarr library"; exit 1; }
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (kept current
|
||||
# every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to a live
|
||||
# fetch when it's stale, and waits out an active rescan before either.
|
||||
SONARR_SERIES_JSON=$(arr_get_tracked_data "sonarr" "$SONARR_URL" "$SONARR_API_KEY" "v3") || { error "Could not fetch Sonarr library"; exit 1; }
|
||||
|
||||
declare -A SONARR_TVDB # tvdb_id → 1
|
||||
declare -A SONARR_TMDB # tmdb_id → 1 (Sonarr v4 exposes tmdbId)
|
||||
@@ -626,9 +630,6 @@ done < <(echo "$SONARR_SERIES_JSON" | jq -r '.[] |
|
||||
|
||||
log "${#SONARR_TVDB[@]} series in Sonarr | ${#EMBY_TMDB_IDS[@]} series in Emby"
|
||||
|
||||
_in_sonarr() { [[ "${SONARR_TVDB["$1"]+x}" || "${SONARR_TMDB["$2"]+x}" ]]; }
|
||||
_in_emby() { [[ "${EMBY_TMDB_IDS["$1"]+x}" || "${EMBY_TVDB_IDS["$2"]+x}" ]]; }
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Score Stage 2 Candidates ━━━
|
||||
# ==============================================================================================
|
||||
Executable
+648
@@ -0,0 +1,648 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================ Radarr Content Classification Scan ==============================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Overseerr lets any user request a movie into the wrong root folder (kids content added to
|
||||
# the general Movies share, anime added to Kids_Movies, etc.) and most users never notice or
|
||||
# correct it. This script reads Radarr's tracked movie list and classifies every movie as
|
||||
# anime / kids-only / regular using metadata signals alone (genre, certification, studio,
|
||||
# original language) — then reports where a movie's computed classification disagrees with
|
||||
# the root folder it's actually sitting in, in both directions:
|
||||
#
|
||||
# FORWARD — a movie classified as anime/kids is sitting outside its dedicated root
|
||||
# REVERSE — a movie sitting inside the kids/anime root doesn't match that classification
|
||||
#
|
||||
# Report-only by default — no files are moved and no Radarr API writes happen unless a
|
||||
# mode flag is given. Pass --move to relocate forward misplacements, or --remove-junk to
|
||||
# delete and import-exclude bad-metadata entries (see OPERATIONAL MODEL below); without
|
||||
# those flags this is purely a detection tool. Every rule below was validated against
|
||||
# this library's real data before being
|
||||
# adopted (see master.conf comments above the curated lists) — this is not a generic
|
||||
# genre-matcher, it's tuned specifically against the false-positive traps that showed up
|
||||
# when testing looser rules (documented per-rule below).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CLASSIFICATION RULES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# is_anime:
|
||||
# (genre Animation AND originalLanguage Japanese) OR studio in RADARR_ANIME_STUDIOS
|
||||
# Always wins over kids when both could apply — explicit priority, not a tiebreak.
|
||||
#
|
||||
# is_kids ("kids will end up watching this alone" — NOT "family movie night"):
|
||||
# not is_anime AND certification not in (R, NC-17) AND (
|
||||
# (genre Animation AND certification != PG-13)
|
||||
# OR studio in RADARR_KIDS_STUDIOS
|
||||
# )
|
||||
# Deliberately excludes bare "Family" genre and bare "G" certification — both genuinely
|
||||
# traced back to live-action films the whole household watches together (Mrs. Doubtfire,
|
||||
# Doctor Dolittle, National Treasure-style adventures, classic Westerns), not kids-only
|
||||
# content. Family movie night stays in the general Movies root by design.
|
||||
#
|
||||
# The PG-13 exclusion on the Animation branch is load-bearing — without it this rule
|
||||
# catches South Park movies, Sausage Party, "9", Resident Evil: Death Island, and (via
|
||||
# the curated studio list) Warner Bros. Animation's R-rated Watchmen films, since that
|
||||
# studio makes both kids content and adult content under the same name.
|
||||
#
|
||||
# is_junk (bad/thin TMDb match, not a real classification problem):
|
||||
# hasFile == false AND imdbId == null AND tmdb votes < RADARR_JUNK_MIN_VOTES
|
||||
# Caught live: two fake "X-Men"/"Wolverine" entries, two fake "Silent Hill" entries, one
|
||||
# fake "The Purge" spinoff — all monitored placeholders with nothing behind them. The
|
||||
# fix for these is removal from Radarr, not blocklist+redownload — there's no release to
|
||||
# blocklist and likely nothing legitimate to redownload under that exact TMDb match.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Report pass (always):
|
||||
# check_api → check_arr_version → arr_get_tracked_data (cache-first, one call)
|
||||
# → classify every movie → report FORWARD, REVERSE and JUNK findings
|
||||
#
|
||||
# Remove-junk pass (--remove-junk, runs first when combined with --move):
|
||||
# One entry at a time, halting on the first failure.
|
||||
# DELETE with deleteFiles=false and addImportExclusion=true — the Radarr entry is
|
||||
# removed and blocked from re-adding, files on disk are never touched. is_junk
|
||||
# requires hasFile == false, so there is no file behind these entries anyway.
|
||||
# Verified by re-fetching and requiring a 404 before counting as removed.
|
||||
#
|
||||
# Move pass (--move):
|
||||
# One movie at a time, verified after each.
|
||||
# hasFile == true → moveFiles=true, then poll the async MoveMovie command to
|
||||
# "completed" (bounded by RADARR_MOVE_POLL_TIMEOUT) before the
|
||||
# DB-field check — the DB flips instantly while the physical
|
||||
# move is still queued.
|
||||
# hasFile == false → correct rootFolderPath/path and trigger MoviesSearch instead;
|
||||
# there is nothing to move.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Report by Default, Act Only on Request
|
||||
# A bare run never calls Radarr's write API and never touches a file — every finding is
|
||||
# just a candidate. Acting on them requires an explicit --move or --remove-junk flag, so
|
||||
# the scan can be scheduled and re-run freely while the curated lists are being tuned
|
||||
# without any risk of it rearranging the library on its own.
|
||||
#
|
||||
# Curated Lists, Not Bare Genre/Cert Matching
|
||||
# Every signal used here failed at least once as a bare/standalone check during rule
|
||||
# development (Family genre, G certification, blanket Animation genre, bare Anime genre
|
||||
# tag, Disney+/general-platform networks) — see master.conf comments for what each
|
||||
# curated list deliberately excludes and why.
|
||||
#
|
||||
# Cache-First, Never a Per-Movie Call
|
||||
# Uses arr_get_tracked_data() same as radarr_cleanup.sh — Radarr's movie list already
|
||||
# embeds everything this script needs per movie, so this is a single API call (or zero,
|
||||
# if the shared cache is warm) regardless of library size.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Required by the container interaction and state writes.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock, plus acquire_lock "wait" around the write passes with an EXIT trap
|
||||
# releasing all locks, so an interrupted run never leaves a lock behind.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases RADARR_URL / RADARR_API_KEY / the root literals.
|
||||
#
|
||||
# curl + jq Dependency Check
|
||||
# Fails fast if either is missing — every classification signal is parsed with jq.
|
||||
#
|
||||
# Report-Only Default
|
||||
# No write happens without --move or --remove-junk.
|
||||
#
|
||||
# Required Var Check
|
||||
# require_var on RADARR_URL and RADARR_API_KEY before any request.
|
||||
#
|
||||
# API Reachability + Version Gate
|
||||
# check_api then check_arr_version against RADARR_VERSION_MAJOR. A major version bump
|
||||
# can move or rename the fields every rule depends on, so a mismatch aborts rather
|
||||
# than classifying against an unknown schema.
|
||||
#
|
||||
# Empty Library Abort
|
||||
# A response of 0 movies aborts — an empty list is indistinguishable from a clean
|
||||
# library and would otherwise report success during an API fault.
|
||||
#
|
||||
# Unconfigured Root Skip
|
||||
# A blank kids/anime root skips that category's checks rather than comparing paths
|
||||
# against an empty string.
|
||||
#
|
||||
# Files Never Deleted
|
||||
# Junk removal passes deleteFiles=false. Only the Radarr entry is removed, and
|
||||
# addImportExclusion=true stops it being re-added. is_junk additionally requires
|
||||
# hasFile == false, so these entries have nothing on disk in the first place.
|
||||
#
|
||||
# One At A Time, Stop On First Failure
|
||||
# Both write passes process one entry at a time and halt on the first failure rather
|
||||
# than continuing through the library.
|
||||
#
|
||||
# Post-Write Verification
|
||||
# Removal is confirmed by re-fetching and requiring a 404. Moves are confirmed by
|
||||
# re-fetching and checking rootFolderPath and hasFile. The API response alone is
|
||||
# never treated as proof.
|
||||
#
|
||||
# Async Move Completion Polling
|
||||
# moveFiles=true flips the DB instantly while the physical move is a separate async
|
||||
# MoveMovie command. Each move polls its own command to "completed" (bounded by
|
||||
# RADARR_MOVE_POLL_TIMEOUT) before the DB-field check, so a batch cannot report
|
||||
# everything moved while files are still queued at the old path.
|
||||
#
|
||||
# Junk Vote Threshold
|
||||
# RADARR_JUNK_MIN_VOTES gates junk detection alongside hasFile == false and a null
|
||||
# imdbId. All three must hold — a thin-metadata entry that actually has a file, or
|
||||
# has an IMDb ID, is never treated as junk.
|
||||
#
|
||||
# Post-Write Cache Refresh
|
||||
# The tracked-data cache is refreshed after writes so no other arr script reads a
|
||||
# stale rootFolderPath or a movie that no longer exists.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
# RADARR_URL / RADARR_API_KEY / RADARR_MOVIES_ROOT — existing, aliased by detect_hosts()
|
||||
# RADARR_KIDS_ROOT / RADARR_ANIME_ROOT — rootFolderPath literals as reported by the API
|
||||
# (e.g. "/kids movies", "/ext-anime-movies") — leave blank on a host with no dedicated
|
||||
# root for that category; the corresponding checks are skipped, not treated as an error.
|
||||
#
|
||||
# master.conf
|
||||
# RADARR_ANIME_STUDIOS / RADARR_KIDS_STUDIOS — curated studio allowlists
|
||||
# RADARR_JUNK_MIN_VOTES — TMDb vote threshold for the bad-metadata check
|
||||
# RADARR_VERSION_MAJOR — expected API major version (reused from radarr_cleanup.sh)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# radarr_classification_scan.sh — normal run, prints report
|
||||
# radarr_classification_scan.sh --log — verbose (per-movie TRACKED-style logging)
|
||||
# radarr_classification_scan.sh --status — show config and exit
|
||||
# radarr_classification_scan.sh --remove-junk — delete + import-exclude bad-metadata entries
|
||||
# radarr_classification_scan.sh --move — relocate forward misplacements (moves the
|
||||
# file if one exists; for hasFile=false
|
||||
# entries, just corrects rootFolderPath/path
|
||||
# and triggers an immediate MoviesSearch)
|
||||
#
|
||||
# --remove-junk and --move can be combined in one run — junk is cleared first, then the
|
||||
# move pass runs against the remaining (now junk-free) classification.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# --move / --remove-junk are script-local flags, not ones parse_args recognizes — check the
|
||||
# raw args before they get filtered into PARSED_ARGS.
|
||||
MOVE_MODE=false
|
||||
REMOVE_JUNK_MODE=false
|
||||
for _arg in "$@"; do
|
||||
[[ "$_arg" == "--move" ]] && MOVE_MODE=true
|
||||
[[ "$_arg" == "--remove-junk" ]] && REMOVE_JUNK_MODE=true
|
||||
done
|
||||
unset _arg
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "${RADARR_URL:-}" ]] || [[ -z "${RADARR_API_KEY:-}" ]]; then
|
||||
info "Radarr not configured on $MY_ID ($LOCAL_SERVER_NAME) — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
require_var RADARR_URL
|
||||
require_var RADARR_API_KEY
|
||||
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Radarr URL: $RADARR_URL"
|
||||
echo "$ICON_GEAR Movies root: $RADARR_MOVIES_ROOT"
|
||||
echo "$ICON_GEAR General root: ${RADARR_GENERAL_ROOT:-<not configured>}"
|
||||
echo "$ICON_GEAR Kids root: ${RADARR_KIDS_ROOT:-<not configured>}"
|
||||
echo "$ICON_GEAR Anime root: ${RADARR_ANIME_ROOT:-<not configured>}"
|
||||
echo "$ICON_GEAR Anime studios: ${#RADARR_ANIME_STUDIOS[@]} curated"
|
||||
echo "$ICON_GEAR Kids studios: ${#RADARR_KIDS_STUDIOS[@]} curated"
|
||||
echo "$ICON_GEAR Junk min votes: ${RADARR_JUNK_MIN_VOTES:-15}"
|
||||
echo "$ICON_GEAR Move mode: $MOVE_MODE"
|
||||
echo "$ICON_GEAR Remove-junk: $REMOVE_JUNK_MODE"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching Radarr Library ━━━"
|
||||
|
||||
if ! check_api "$RADARR_URL" "Radarr" 10; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
check_arr_version "$RADARR_URL" "$RADARR_API_KEY" "v3" "$RADARR_VERSION_MAJOR" "Radarr" || exit 1
|
||||
|
||||
MOVIES_RESPONSE=$(arr_get_tracked_data "radarr" "$RADARR_URL" "$RADARR_API_KEY" "v3") || {
|
||||
error "Failed to fetch movies from Radarr"
|
||||
exit 1
|
||||
}
|
||||
|
||||
MOVIE_COUNT=$(echo "$MOVIES_RESPONSE" | jq -r 'length' 2>/dev/null)
|
||||
if [[ -z "$MOVIE_COUNT" ]] || [[ "$MOVIE_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 movies — aborting"
|
||||
exit 1
|
||||
fi
|
||||
info "$MOVIE_COUNT movies loaded"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Classify ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Classifying ━━━"
|
||||
|
||||
ANIME_STUDIOS_JSON=$(printf '%s\n' "${RADARR_ANIME_STUDIOS[@]}" | jq -R . | jq -s .)
|
||||
KIDS_STUDIOS_JSON=$(printf '%s\n' "${RADARR_KIDS_STUDIOS[@]}" | jq -R . | jq -s .)
|
||||
JUNK_MIN_VOTES="${RADARR_JUNK_MIN_VOTES:-15}"
|
||||
|
||||
RESULTS=$(echo "$MOVIES_RESPONSE" | jq \
|
||||
--argjson animeStudios "$ANIME_STUDIOS_JSON" \
|
||||
--argjson kidsStudios "$KIDS_STUDIOS_JSON" \
|
||||
--arg animeRoot "${RADARR_ANIME_ROOT:-}" \
|
||||
--arg kidsRoot "${RADARR_KIDS_ROOT:-}" \
|
||||
--argjson junkMinVotes "$JUNK_MIN_VOTES" '
|
||||
def is_anime:
|
||||
(any(.genres[]?; . == "Animation") and .originalLanguage.name == "Japanese")
|
||||
or (.studio as $s | $animeStudios | index($s) != null);
|
||||
def not_adult: (.certification != "R") and (.certification != "NC-17");
|
||||
def is_kids:
|
||||
(is_anime | not) and not_adult and (
|
||||
(any(.genres[]?; . == "Animation") and .certification != "PG-13")
|
||||
or (.studio as $s | $kidsStudios | index($s) != null)
|
||||
);
|
||||
def is_junk:
|
||||
(.hasFile == false) and (.imdbId == null)
|
||||
and ((.ratings.tmdb.votes // 999999) < $junkMinVotes);
|
||||
|
||||
map(
|
||||
{
|
||||
title, id, studio, certification, rootFolderPath, genres, hasFile,
|
||||
is_anime: is_anime,
|
||||
is_kids: is_kids,
|
||||
is_junk: is_junk
|
||||
} |
|
||||
. + {
|
||||
forward_anime_miss: (.is_anime and $animeRoot != "" and .rootFolderPath != $animeRoot),
|
||||
forward_kids_miss: (.is_kids and $kidsRoot != "" and .rootFolderPath != $kidsRoot),
|
||||
reverse_anime_leak: ((.is_anime | not) and $animeRoot != "" and .rootFolderPath == $animeRoot and (.is_junk | not)),
|
||||
reverse_kids_leak: ((.is_anime | not) and (.is_kids | not) and $kidsRoot != "" and .rootFolderPath == $kidsRoot
|
||||
and (.is_junk | not)
|
||||
and (.certification == "R" or .certification == "NC-17"
|
||||
or (.certification == "PG-13"
|
||||
and (any(.genres[]?; . == "Family") | not)
|
||||
and (any(.genres[]?; . == "Animation") | not))))
|
||||
}
|
||||
)
|
||||
')
|
||||
|
||||
FORWARD_ANIME_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.forward_anime_miss)] | length')
|
||||
FORWARD_KIDS_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.forward_kids_miss)] | length')
|
||||
REVERSE_ANIME_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.reverse_anime_leak)] | length')
|
||||
REVERSE_KIDS_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.reverse_kids_leak)] | length')
|
||||
JUNK_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.is_junk)] | length')
|
||||
|
||||
if [[ "$ENABLE_LOGGING" == true ]]; then
|
||||
echo "$RESULTS" | jq -r '.[] | select(.forward_anime_miss or .forward_kids_miss or .reverse_anime_leak or .reverse_kids_leak or .is_junk) |
|
||||
" [\(if .is_junk then "JUNK" elif .forward_anime_miss then "FORWARD-ANIME" elif .forward_kids_miss then "FORWARD-KIDS" elif .reverse_anime_leak then "REVERSE-ANIME" elif .reverse_kids_leak then "REVERSE-KIDS" else "?" end)] \(.title) (root: \(.rootFolderPath), studio: \(.studio // "n/a"), cert: \(.certification // "n/a"))"'
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY RADARR CLASSIFICATION SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Movies scanned: $MOVIE_COUNT"
|
||||
echo "$ICON_TRASH Forward — anime miss: $FORWARD_ANIME_COUNT (classified anime, outside ${RADARR_ANIME_ROOT:-<unconfigured>})"
|
||||
echo "$ICON_TRASH Forward — kids miss: $FORWARD_KIDS_COUNT (classified kids, outside ${RADARR_KIDS_ROOT:-<unconfigured>})"
|
||||
echo "$ICON_WARN Reverse — anime leak: $REVERSE_ANIME_COUNT (in ${RADARR_ANIME_ROOT:-<unconfigured>}, no anime signal — review, may be deliberate style placement)"
|
||||
echo "$ICON_WARN Reverse — kids leak: $REVERSE_KIDS_COUNT (in ${RADARR_KIDS_ROOT:-<unconfigured>}, adult-rated content)"
|
||||
echo "$ICON_PROTECTED Bad metadata (junk): $JUNK_COUNT (hasFile=false, no imdbId, thin TMDb match — candidates for removal, not redownload)"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
[[ "$ENABLE_LOGGING" != true ]] && echo " (run with --log for the per-title list)"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Remove-Junk Mode ━━━
|
||||
# ==============================================================================================
|
||||
# Junk entries are bad/thin TMDb matches with hasFile=false — there's no release to blocklist
|
||||
# (nothing was ever grabbed) and no file to delete, only a bad monitored record. The fix is
|
||||
# removing the record and adding it to Radarr's import exclusion list (same mechanism
|
||||
# radarr_tmdb_removed.sh already uses via RADARR_DROPPED_ADD_EXCLUSION) so the same bad TMDb
|
||||
# match can't get re-added by a future Overseerr request or list sync.
|
||||
if [[ "$REMOVE_JUNK_MODE" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_TRASH Remove-Junk Mode ━━━"
|
||||
|
||||
acquire_lock "wait"
|
||||
trap "_release_all_locks" EXIT
|
||||
|
||||
JUNK_TARGETS=$(echo "$RESULTS" | jq -c '[.[] | select(.is_junk)]')
|
||||
JUNK_TARGET_COUNT=$(echo "$JUNK_TARGETS" | jq 'length')
|
||||
|
||||
if [[ "$JUNK_TARGET_COUNT" -eq 0 ]]; then
|
||||
info "No junk entries to remove"
|
||||
else
|
||||
warn "About to remove $JUNK_TARGET_COUNT junk entries — one at a time, verifying after each"
|
||||
|
||||
JUNK_REMOVED=0
|
||||
JUNK_FAILED=0
|
||||
|
||||
while IFS= read -r item; do
|
||||
id=$(echo "$item" | jq -r '.id')
|
||||
title=$(echo "$item" | jq -r '.title')
|
||||
|
||||
http_code=$(curl -sf -o /dev/null -w "%{http_code}" -X DELETE \
|
||||
--max-time 15 \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
"${RADARR_URL}/api/v3/movie/${id}?deleteFiles=false&addImportExclusion=true" 2>/dev/null)
|
||||
|
||||
if [[ "$http_code" != "200" && "$http_code" != "202" ]]; then
|
||||
error " ✗ $title — API returned HTTP $http_code — stopping (review before re-running)"
|
||||
(( JUNK_FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
sleep 1
|
||||
|
||||
# Verify — the movie should now be gone entirely (404).
|
||||
verify_code=$(curl -sf -o /dev/null -w "%{http_code}" \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
"${RADARR_URL}/api/v3/movie/${id}" 2>/dev/null)
|
||||
|
||||
if [[ "$verify_code" == "404" ]]; then
|
||||
echo " $ICON_SUCCESS $title — removed and excluded"
|
||||
(( JUNK_REMOVED++ ))
|
||||
else
|
||||
error " ✗ $title — verification failed (still returns HTTP $verify_code) — stopping"
|
||||
(( JUNK_FAILED++ ))
|
||||
break
|
||||
fi
|
||||
done < <(echo "$JUNK_TARGETS" | jq -c '.[]')
|
||||
|
||||
if [[ "$JUNK_REMOVED" -gt 0 ]]; then
|
||||
info "Refreshing shared tracked-data cache..."
|
||||
fresh_movies=$(arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "movie" "Radarr")
|
||||
[[ -n "$fresh_movies" ]] && arr_cache_write "radarr" "$fresh_movies"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY REMOVE-JUNK SUMMARY ━━━━━"
|
||||
echo "$ICON_SUCCESS Removed: $JUNK_REMOVED"
|
||||
echo "$ICON_ERROR Failed: $JUNK_FAILED"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
fi
|
||||
|
||||
_release_all_locks
|
||||
trap - EXIT
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Move Mode ━━━
|
||||
# ==============================================================================================
|
||||
# Acts on FORWARD misplacements (clear-cut: classified anime/kids, sitting in the wrong root)
|
||||
# and on REVERSE-KIDS leaks (adult certification with zero Family/Animation genre sitting in
|
||||
# the kids root — also clear-cut, moved back to RADARR_GENERAL_ROOT). Does NOT act on
|
||||
# REVERSE-ANIME leaks — those are genuine judgment calls, since deliberate style placements
|
||||
# like Castlevania/Legend of Korra legitimately live in the anime root without matching the
|
||||
# anime signal — or on JUNK (those need removal from Radarr, not a file move).
|
||||
#
|
||||
# One movie at a time, verified after each. moveFiles=true flips the DB (rootFolderPath/
|
||||
# hasFile) instantly, but the physical move is a separate async MoveMovie command Radarr's
|
||||
# own MoveMovieService drains one at a time internally — same architecture that raced on the
|
||||
# Sonarr side (episodeFileCount reported at the new path via API while the real files were
|
||||
# still sitting at the old one, MoveSeries command queued behind ~20 others). Never confirmed
|
||||
# live on the Radarr side, but the same DB-write-is-instant/move-is-async split applies, so
|
||||
# each move here polls its own MoveMovie command to "completed" before the DB-field check
|
||||
# runs, mirroring the Sonarr fix.
|
||||
if [[ "$MOVE_MODE" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Move Mode — Forward Misplacements ━━━"
|
||||
|
||||
acquire_lock "wait"
|
||||
trap "_release_all_locks" EXIT
|
||||
|
||||
build_arr_path_map "RADARR"
|
||||
|
||||
# hasFile==false entries (monitored but never downloaded) have nothing to physically
|
||||
# move — confirmed live: "Biohazard 4: Incubate" had hasFile=false even in the cache from
|
||||
# before any of this ran, not something this script broke. There's still real value in
|
||||
# fixing them, though: correct the DB pointer now (so Radarr saves to the right root
|
||||
# whenever it does find a release) and kick off an immediate search rather than waiting
|
||||
# for the next scheduled one. Junk entries are still excluded entirely — nothing to
|
||||
# search for there, they need removal instead.
|
||||
#
|
||||
# reverse_kids_leak is included here (unlike reverse_anime_leak) because its signal is
|
||||
# specifically "adult certification with zero Family/Animation genre" — there's no
|
||||
# legitimate stylistic reason for that combination to sit in a curated kids library, unlike
|
||||
# the anime side where deliberate style placements (Castlevania, Legend of Korra) are
|
||||
# common and valid. Confirmed live: the 4 titles this caught (The Addams Family, Saving
|
||||
# Mr. Banks, Dark Shadows, The DUFF) are all genuinely non-kids content, not judgment calls.
|
||||
MOVE_TARGETS=$(echo "$RESULTS" | jq -c '[.[] | select((.forward_anime_miss or .forward_kids_miss or .reverse_kids_leak) and (.is_junk | not))]')
|
||||
MOVE_COUNT=$(echo "$MOVE_TARGETS" | jq 'length')
|
||||
|
||||
if [[ "$MOVE_COUNT" -eq 0 ]]; then
|
||||
info "Nothing to move"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
warn "About to process $MOVE_COUNT movies — one at a time, verifying after each"
|
||||
|
||||
MOVED=0
|
||||
RELOCATED_SEARCH=0
|
||||
FAILED=0
|
||||
|
||||
while IFS= read -r item; do
|
||||
id=$(echo "$item" | jq -r '.id')
|
||||
title=$(echo "$item" | jq -r '.title')
|
||||
is_anime_flag=$(echo "$item" | jq -r '.is_anime')
|
||||
is_forward_kids=$(echo "$item" | jq -r '.forward_kids_miss')
|
||||
had_file=$(echo "$item" | jq -r '.hasFile')
|
||||
if [[ "$is_anime_flag" == "true" ]]; then
|
||||
target_root="$RADARR_ANIME_ROOT"
|
||||
elif [[ "$is_forward_kids" == "true" ]]; then
|
||||
target_root="$RADARR_KIDS_ROOT"
|
||||
else
|
||||
target_root="$RADARR_GENERAL_ROOT"
|
||||
fi
|
||||
|
||||
if [[ -z "$target_root" ]]; then
|
||||
error " ✗ $title — target root not configured (RADARR_GENERAL_ROOT blank), skipping"
|
||||
(( FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# RESULTS only carries the reduced report fields — Radarr's PUT expects the complete
|
||||
# resource representation, so fetch a fresh full movie record to modify and send back.
|
||||
full_movie=$(arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "movie/$id" "Radarr")
|
||||
if [[ -z "$full_movie" ]]; then
|
||||
error " ✗ $title — could not fetch full movie record, skipping"
|
||||
(( FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
old_path=$(echo "$full_movie" | jq -r '.path')
|
||||
folder_name="${old_path##*/}"
|
||||
|
||||
# A literal "/" in the folder name would build a broken nested directory instead of
|
||||
# moving to one clean folder — bit us once already doing this by hand for Sonarr.
|
||||
if [[ "$folder_name" == *"/"* ]]; then
|
||||
error " ✗ $title — folder name contains '/', skipping (needs manual handling)"
|
||||
(( FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
new_path="${target_root}/${folder_name}"
|
||||
|
||||
if [[ "$had_file" == "true" ]]; then
|
||||
info " → $title: $old_path → $new_path (moving file)"
|
||||
move_qs="?moveFiles=true"
|
||||
else
|
||||
info " → $title: $old_path → $new_path (no file — relocating + search)"
|
||||
move_qs=""
|
||||
fi
|
||||
|
||||
updated_movie=$(echo "$full_movie" | jq --arg root "$target_root" --arg path "$new_path" \
|
||||
'.rootFolderPath = $root | .path = $path')
|
||||
|
||||
http_code=$(curl -sf -o /dev/null -w "%{http_code}" -X PUT \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$updated_movie" \
|
||||
"${RADARR_URL}/api/v3/movie/${id}${move_qs}" 2>/dev/null)
|
||||
|
||||
if [[ "$http_code" != "200" && "$http_code" != "202" ]]; then
|
||||
error " ✗ $title — API returned HTTP $http_code — stopping (review before re-running)"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
# moveFiles=true flips rootFolderPath/hasFile in the DB instantly, but the actual
|
||||
# physical move is a separate async MoveMovie command that Radarr's MoveMovieService
|
||||
# drains one at a time internally — mirrors the confirmed Sonarr race (see header
|
||||
# comment above Move Mode). Poll the actual command to completion before trusting the
|
||||
# DB-field check below.
|
||||
if [[ -n "$move_qs" ]]; then
|
||||
move_cmd_id=""
|
||||
for _ in 1 2 3 4 5; do
|
||||
move_cmd_id=$(curl -sf --max-time 10 -H "X-Api-Key: $RADARR_API_KEY" \
|
||||
"${RADARR_URL}/api/v3/command" 2>/dev/null | \
|
||||
jq -r --argjson mid "$id" \
|
||||
'[.[] | select(.name == "MoveMovie" and .body.movieId == $mid)] | sort_by(.id) | last | .id // empty' \
|
||||
2>/dev/null)
|
||||
[[ -n "$move_cmd_id" ]] && break
|
||||
sleep 1
|
||||
done
|
||||
|
||||
if [[ -z "$move_cmd_id" ]]; then
|
||||
error " ✗ $title — could not locate the MoveMovie command — stopping (review before re-running)"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
info " → $title: MoveMovie command $move_cmd_id queued, waiting for completion..."
|
||||
move_status="" move_polled=0
|
||||
while [[ "$move_polled" -lt "$RADARR_MOVE_POLL_TIMEOUT" ]]; do
|
||||
move_status=$(curl -sf --max-time 10 -H "X-Api-Key: $RADARR_API_KEY" \
|
||||
"${RADARR_URL}/api/v3/command/${move_cmd_id}" 2>/dev/null | \
|
||||
jq -r '.status // empty' 2>/dev/null)
|
||||
[[ "$move_status" == "completed" || "$move_status" == "failed" ]] && break
|
||||
sleep 10
|
||||
(( move_polled += 10 ))
|
||||
[[ $(( move_polled % 60 )) -eq 0 ]] && log " still moving $title... (${move_polled}s elapsed)"
|
||||
done
|
||||
|
||||
if [[ "$move_status" != "completed" ]]; then
|
||||
error " ✗ $title — MoveMovie command $move_cmd_id ended as '${move_status:-timed out after ${RADARR_MOVE_POLL_TIMEOUT}s}' — stopping"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
fi
|
||||
|
||||
sleep 3
|
||||
|
||||
# Never trust the PUT response alone — re-fetch and confirm the change actually landed.
|
||||
verify_movie=$(arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "movie/$id" "Radarr")
|
||||
verify_root=$(echo "$verify_movie" | jq -r '.rootFolderPath')
|
||||
verify_hasfile=$(echo "$verify_movie" | jq -r '.hasFile')
|
||||
|
||||
if [[ "$verify_root" != "$target_root" ]]; then
|
||||
error " ✗ $title — verification failed (root: $verify_root) — stopping"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
if [[ "$had_file" == "true" ]]; then
|
||||
if [[ "$verify_hasfile" == "true" ]]; then
|
||||
echo " $ICON_SUCCESS $title — moved and verified"
|
||||
(( MOVED++ ))
|
||||
else
|
||||
error " ✗ $title — verification failed (root updated but hasFile now false) — stopping"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
else
|
||||
search_code=$(curl -sf -o /dev/null -w "%{http_code}" -X POST \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{\"name\":\"MoviesSearch\",\"movieIds\":[${id}]}" \
|
||||
"${RADARR_URL}/api/v3/command" 2>/dev/null)
|
||||
if [[ "$search_code" == "200" || "$search_code" == "201" ]]; then
|
||||
echo " $ICON_SUCCESS $title — relocated, search triggered"
|
||||
else
|
||||
warn " $title — relocated but search trigger returned HTTP $search_code (will pick up on next scheduled search)"
|
||||
fi
|
||||
(( RELOCATED_SEARCH++ ))
|
||||
fi
|
||||
done < <(echo "$MOVE_TARGETS" | jq -c '.[]')
|
||||
|
||||
# arr_get_tracked_data() is cache-first — every write above changed rootFolderPath, so the
|
||||
# shared cache is now stale until the next scheduled arr_cache_prefill run (up to 30min).
|
||||
# Every other script reading this cache (cleanup, discovery, etc.) would see wrong data
|
||||
# until then — refresh it now with one more live fetch rather than leave that window open.
|
||||
if [[ "$(( MOVED + RELOCATED_SEARCH ))" -gt 0 ]]; then
|
||||
info "Refreshing shared tracked-data cache..."
|
||||
fresh_movies=$(arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "movie" "Radarr")
|
||||
[[ -n "$fresh_movies" ]] && arr_cache_write "radarr" "$fresh_movies"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY MOVE SUMMARY ━━━━━"
|
||||
echo "$ICON_SUCCESS Moved (file relocated): $MOVED"
|
||||
echo "$ICON_SUCCESS Relocated + search triggered: $RELOCATED_SEARCH"
|
||||
echo "$ICON_ERROR Failed: $FAILED"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
fi
|
||||
|
||||
exit 0
|
||||
@@ -0,0 +1,696 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Radarr Cleanup =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Delete orphaned movie files not tracked by Radarr. Queries the API for all
|
||||
# tracked movie file paths, walks the library on disk, and removes anything
|
||||
# untracked that is old enough to be past the import window. Triggers an Emby
|
||||
# library clean after each deletion run so ghost entries disappear immediately.
|
||||
#
|
||||
# The tracked-count floor check (Safety Layer 6) is rescan-aware: if Radarr's own
|
||||
# RescanMovie/DownloadedMoviesScan is active (independently of this script's own
|
||||
# lighter ProcessMonitoredDownloads pre-flight), a genuinely low mid-scan count gets
|
||||
# waited out (calibrated to that command's historical duration via
|
||||
# arr_get_rescan_duration(), up to 3 strikes) and re-fetched rather than triggering a
|
||||
# false-alarm abort. Mirrors the same fix built for lidarr_cleanup.sh 2026-07-16 after
|
||||
# a whole-library rescan there made trackFileCount read 22% of normal mid-scan.
|
||||
#
|
||||
# Cache-first movie list (2026-07-17), batched moviefile fetch (2026-07-19). The movie list
|
||||
# comes from the shared tracked-data cache via arr_get_tracked_data() — fresh (kept warm
|
||||
# every 30min by arr_cache_prefill.sh), live fetch as fallback. Radarr's movie list embeds
|
||||
# movieFile.path directly on every hasFile=true entry, but that's only the *primary* file —
|
||||
# Radarr 6+ supports a second tracked file per movie (alternate editions/extras) that never
|
||||
# shows up there, so relying on it alone misclassified a movie's second edition as an orphan
|
||||
# (confirmed live 2026-07-19: The Crash, They Will Kill You, The Drama, Lee Cronin's The
|
||||
# Mummy, and Ready or Not: Here I Come all had a legitimately-tracked second file deleted-flagged
|
||||
# this way). /moviefile?movieId=X returns every file for a movie, including secondaries, and
|
||||
# accepts movieId as a repeated query param for a bulk fetch — but the whole library in one
|
||||
# request 414s (Request-URI Too Long, confirmed live), so _fetch_tracked_files() batches
|
||||
# movieId params BATCH_SIZE at a time instead: ~14 requests for a ~2800-movie library rather
|
||||
# than the up-to-2896 individual per-movie calls the 2026-07-17 optimization eliminated, and
|
||||
# rather than the one-shot list read that missed secondary files. The filesystem is walked
|
||||
# once per run, not twice — classification records which paths are eligible for deletion as
|
||||
# it goes, and the delete pass (once the size-threshold check below passes) just acts on that
|
||||
# list instead of re-walking and re-classifying the whole tree. That single walk also gets
|
||||
# size+ctime straight from find -printf instead of a separate stat fork per file — find
|
||||
# already has to stat() every entry to know it's -type f, so this is free by comparison.
|
||||
# Measured ~130x faster per file (0.033ms vs 4.3ms).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Every file encountered on disk is classified into one of five categories:
|
||||
#
|
||||
# TRACKED — Radarr API knows this exact path → leave it alone
|
||||
# PROTECTED — matches RADARR_PROTECTED_PATTERNS → never delete
|
||||
# ORPHAN — video file, not tracked, older than RADARR_ORPHAN_AGE → delete
|
||||
# JUNK — not a video extension, not protected → delete regardless of age
|
||||
# RECENT — not tracked, under RADARR_ORPHAN_AGE → skip (may be mid-import)
|
||||
#
|
||||
# Radarr generates movie artwork (*.jpg), metadata (*.nfo), and manages subtitles
|
||||
# (*.srt, *.sub, *.ass) but does NOT include these in its tracked file API response.
|
||||
# Without PROTECTED classification these would be deleted — breaking Radarr and
|
||||
# Emby metadata display.
|
||||
#
|
||||
# After deletions: notify_emby_scan() triggers Emby "Clean Missing Files" task.
|
||||
# Emby removes ghost entries immediately — no user-facing file-not-found errors.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# API as Ground Truth
|
||||
# What Radarr tracks is authoritative. Files not in the API response are
|
||||
# orphans — Radarr has no record of them and they serve no purpose.
|
||||
# The script never infers ownership from directory structure alone.
|
||||
#
|
||||
# Age Gate Before Deletion
|
||||
# Files under RADARR_ORPHAN_AGE are left alone regardless of tracked status.
|
||||
# Radarr's import pipeline writes files before registering them — acting
|
||||
# immediately would delete files mid-import.
|
||||
# Age is measured from ctime, not mtime — an import preserves the release's original
|
||||
# mtime, so a file that landed today can read as years old and skip this gate. Depends
|
||||
# on media_shares_permissions.sh touching only entries that are actually wrong.
|
||||
#
|
||||
# Emby Cleanup Is Part of the Job
|
||||
# Deleting a file without telling Emby leaves ghost entries that show as
|
||||
# broken items. Triggering the Emby clean is not optional — it completes
|
||||
# the deletion from the user's perspective.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Seven gates — ALL must pass before any file is touched:
|
||||
# 1. Container running and not starting/unhealthy
|
||||
# 2. API reachable
|
||||
# 3. API version matches RADARR_VERSION_MAJOR in master.conf
|
||||
# 4. Movie count > 0
|
||||
# 5. Tracked file count > 0
|
||||
# 6. Tracked count >= RADARR_MIN_TRACKED_PCT % of last known count
|
||||
# 7. Deletion size < RADARR_MAX_DELETE_GB — or --i-know-what-im-doing required
|
||||
#
|
||||
# acquire_lock "wait" — large scans take time, wait for previous run to finish
|
||||
# jq + curl validation — exits if either tool missing
|
||||
# ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT)
|
||||
# notify_emby_scan() — triggers Emby clean after deletion
|
||||
# Silent by default — orphans/junk warn(), clean library logs silently
|
||||
#
|
||||
# ==============================================================================================
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# RADARR_TRACKED_COUNT_FILE — persistent baseline for the tracked % safety check (gate 6)
|
||||
# Updated after each successful run. Protects against misconfigured root path
|
||||
# returning an empty API response and deleting the entire library.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_RADARR_URL / HOST*_RADARR_API_KEY / HOST*_RADARR_MOVIES_ROOT
|
||||
# HOST*_RADARR_PATH_MAP — container path → host path translation
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# RADARR_ORPHAN_AGE — days before untracked file eligible for deletion
|
||||
# RADARR_MAX_DELETE_GB — require --i-know-what-im-doing above this
|
||||
# RADARR_MIN_TRACKED_PCT — abort if tracked count drops below this % of last run
|
||||
# RADARR_TRACKED_COUNT_FILE — persistent baseline file path
|
||||
# RADARR_EXTENSIONS — video file extensions for orphan classification
|
||||
# RADARR_PROTECTED_PATTERNS — file patterns never deleted
|
||||
# RADARR_VERSION_MAJOR — expected Radarr major version for API safety check
|
||||
# RADARR_IMPORT_SCAN_TIMEOUT — seconds to wait for pre-flight import scan (default 600)
|
||||
# ARR_CLEANUP_STATS — stats file path (read by coffee report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# radarr_cleanup.sh — normal run
|
||||
# radarr_cleanup.sh --dry-run — preview, no deletions
|
||||
# radarr_cleanup.sh --log — verbose output
|
||||
# radarr_cleanup.sh --status — show config and exit
|
||||
# radarr_cleanup.sh --i-know-what-im-doing — bypass size threshold
|
||||
# radarr_cleanup.sh --i-know-what-im-doing --skip-age-check — NUCLEAR MODE
|
||||
#
|
||||
# NUCLEAR MODE: both flags bypass age check AND size threshold. User accepts full
|
||||
# responsibility — the flag name is long and annoying by design.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# ── Special flag pre-processing ───────────────────────────────────────────────────────────────
|
||||
parse_destructive_flags "$@"
|
||||
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
|
||||
# ── Nuclear mode warning ──────────────────────────────────────────────────────────────────────
|
||||
nuclear_mode_warning
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for Radarr API calls"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
notify "Radarr cleanup failed on $(hostname) — jq not installed" "Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
||||
acquire_lock "wait"
|
||||
TMP_DIR="/tmp/radarr_cleanup_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "_release_all_locks; rm -rf $TMP_DIR" EXIT
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases RADARR_URL, RADARR_API_KEY, RADARR_MOVIES_ROOT
|
||||
detect_hosts
|
||||
|
||||
# Skip if Radarr is not configured on this host
|
||||
if [[ -z "${RADARR_URL:-}" ]] || [[ -z "${RADARR_API_KEY:-}" ]]; then
|
||||
info "Radarr not configured on $MY_ID ($LOCAL_SERVER_NAME) — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ARR_DOCKER_TIMEOUT=15
|
||||
RADARR_CONTAINER="Radarr"
|
||||
|
||||
# Build path map from MY_ID's Radarr path map
|
||||
build_arr_path_map "RADARR"
|
||||
|
||||
require_var RADARR_URL
|
||||
require_var RADARR_API_KEY
|
||||
require_var RADARR_MOVIES_ROOT
|
||||
|
||||
if [[ ! -d "$RADARR_MOVIES_ROOT" ]]; then
|
||||
error "Movies root not found: $RADARR_MOVIES_ROOT"
|
||||
notify "Radarr cleanup failed on $(hostname) — movies root not found: $RADARR_MOVIES_ROOT" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Config: url=${RADARR_URL} root=${RADARR_MOVIES_ROOT}"
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $RADARR_URL"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be deleted"
|
||||
[[ "$I_KNOW" == true ]] && warn "OVERRIDE — --i-know-what-im-doing active"
|
||||
[[ "$SKIP_AGE_CHECK" == true ]] && warn "OVERRIDE — --skip-age-check active — age check bypassed"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Radarr URL: $RADARR_URL"
|
||||
echo "$ICON_GEAR Movies root: $RADARR_MOVIES_ROOT"
|
||||
echo "$ICON_TIME Orphan age: ${RADARR_ORPHAN_AGE} days"
|
||||
echo "$ICON_GEAR Max delete: ${RADARR_MAX_DELETE_GB}GB (requires --i-know-what-im-doing)"
|
||||
echo "$ICON_GEAR Min tracked %: ${RADARR_MIN_TRACKED_PCT}%"
|
||||
echo "$ICON_GEAR Radarr ver: v${RADARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Extensions: ${RADARR_EXTENSIONS[*]}"
|
||||
echo "$ICON_GEAR Protected patterns: ${RADARR_PROTECTED_PATTERNS[*]}"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "$ICON_GEAR I know: $I_KNOW"
|
||||
echo "$ICON_GEAR Skip age check: $SKIP_AGE_CHECK"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 1 — Container Health ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
|
||||
check_container_health "$RADARR_CONTAINER" "$ARR_DOCKER_TIMEOUT" "Radarr Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# check_container_health(), arr_api(), has_extension(), matches_pattern_list(), format_bytes() — common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Pre-flight: Radarr Import Scan ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Pre-flight: Radarr Import Scan ━━━"
|
||||
|
||||
# Fetch root folders from Radarr API and translate container paths to host paths
|
||||
mapfile -t SCAN_ROOTS < <(
|
||||
arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "rootfolder" "Radarr" | \
|
||||
jq -r '.[].path' 2>/dev/null | \
|
||||
while IFS= read -r cp; do translate_path "$cp"; done
|
||||
)
|
||||
|
||||
if [[ "${#SCAN_ROOTS[@]}" -eq 0 ]]; then
|
||||
error "No root folders returned from Radarr API — aborting"
|
||||
notify "Radarr cleanup aborted on $(hostname) — no root folders from API" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "Scan targets (${#SCAN_ROOTS[@]}): ${SCAN_ROOTS[*]}"
|
||||
|
||||
info "Triggering ProcessMonitoredDownloads pre-flight"
|
||||
SCAN_PAYLOAD='{"name": "ProcessMonitoredDownloads"}'
|
||||
|
||||
trigger_and_await_command "$RADARR_URL" "$RADARR_API_KEY" "v3" "$SCAN_PAYLOAD" "${RADARR_IMPORT_SCAN_TIMEOUT:-600}" "radarr"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Radarr Tracked Files ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching Radarr Tracked Files ━━━"
|
||||
|
||||
# Safety Layer 2 — API reachability
|
||||
if ! check_api "$RADARR_URL" "Radarr" 10; then
|
||||
notify "Radarr cleanup aborted on $(hostname) — API unreachable" "Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Safety Layer 3 — API version check
|
||||
check_arr_version "$RADARR_URL" "$RADARR_API_KEY" "v3" "$RADARR_VERSION_MAJOR" "Radarr" || exit 1
|
||||
|
||||
info "Querying Radarr API..."
|
||||
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (now kept
|
||||
# current every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to
|
||||
# a live fetch when it's stale, and waits out an active rescan before either. Only the movie
|
||||
# list itself is cached — the per-movie moviefile data below is never cached and always live,
|
||||
# since that's the actual disk-truth this script's cleanup decisions depend on.
|
||||
MOVIES_RESPONSE=$(arr_get_tracked_data "radarr" "$RADARR_URL" "$RADARR_API_KEY" "v3") || {
|
||||
error "Failed to fetch movies from Radarr"
|
||||
notify "Radarr cleanup failed on $(hostname) — could not fetch movies" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
}
|
||||
|
||||
MOVIE_IDS=$(echo "$MOVIES_RESPONSE" | jq -r '.[].id' 2>/dev/null)
|
||||
MOVIE_COUNT=$(echo "$MOVIE_IDS" | grep -c "." 2>/dev/null || true)
|
||||
|
||||
# Safety Layer 4 — movie count > 0
|
||||
if [[ "$MOVIE_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 movies — aborting to prevent mass deletion"
|
||||
notify "Radarr cleanup aborted on $(hostname) — 0 movies returned" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "Found $MOVIE_COUNT movies — fetching movie files..."
|
||||
|
||||
TRACKED_FILE="$TMP_DIR/tracked_paths.txt"
|
||||
> "$TRACKED_FILE"
|
||||
|
||||
# Fetches every movie's file path(s) fresh into TRACKED_FILE/TRACKED_MAP/TRACKED_COUNT.
|
||||
# Pulled into a function so the rescan-aware retry below can re-fetch after waiting without
|
||||
# duplicating this whole loop inline.
|
||||
#
|
||||
# Batched, not per-movie and not a single one-shot list read (2026-07-19) — see the header
|
||||
# comment above for why movie.movieFile.path alone misses secondary edition files. Still a
|
||||
# fresh live fetch on every call (not cache-first) — this function's whole purpose during the
|
||||
# rescan-aware retry below is to see Radarr's progress as the rescan updates hasFile/
|
||||
# movieFile, so it needs genuinely current data each time, not a stale snapshot.
|
||||
_fetch_tracked_files() {
|
||||
> "$TRACKED_FILE"
|
||||
local movies_now
|
||||
movies_now=$(arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "movie" "Radarr" 2>/dev/null)
|
||||
|
||||
local _ids=() _id _qs="" _batch_count=0
|
||||
local BATCH_SIZE=200 # 250 confirmed working live 2026-07-19; kept under that for margin
|
||||
mapfile -t _ids < <(echo "$movies_now" | jq -r '.[] | select(.hasFile==true) | .id' 2>/dev/null)
|
||||
|
||||
{
|
||||
for _id in "${_ids[@]}"; do
|
||||
_qs+="movieId=${_id}&"
|
||||
(( _batch_count++ ))
|
||||
if [[ "$_batch_count" -ge "$BATCH_SIZE" ]]; then
|
||||
arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "moviefile?${_qs%&}" "Radarr" 2>/dev/null | \
|
||||
jq -r '.[].path' 2>/dev/null
|
||||
_qs=""
|
||||
_batch_count=0
|
||||
fi
|
||||
done
|
||||
if [[ -n "$_qs" ]]; then
|
||||
arr_api "$RADARR_URL" "$RADARR_API_KEY" "v3" "moviefile?${_qs%&}" "Radarr" 2>/dev/null | \
|
||||
jq -r '.[].path' 2>/dev/null
|
||||
fi
|
||||
} | while IFS= read -r api_path; do
|
||||
[[ -z "$api_path" ]] && continue
|
||||
translate_path "$api_path" >> "$TRACKED_FILE"
|
||||
done
|
||||
|
||||
sort -u "$TRACKED_FILE" -o "$TRACKED_FILE"
|
||||
|
||||
# Build in-memory lookup map — O(1) per lookup vs O(n) grep per file
|
||||
# Eliminates the main performance bottleneck for large libraries
|
||||
unset TRACKED_MAP
|
||||
declare -gA TRACKED_MAP
|
||||
while IFS= read -r _tracked_path; do
|
||||
[[ -n "$_tracked_path" ]] && TRACKED_MAP["$_tracked_path"]=1
|
||||
done < "$TRACKED_FILE"
|
||||
unset _tracked_path
|
||||
TRACKED_COUNT=$(wc -l < "$TRACKED_FILE")
|
||||
}
|
||||
|
||||
_fetch_tracked_files
|
||||
info "Built in-memory lookup map: ${#TRACKED_MAP[@]} tracked paths"
|
||||
|
||||
# Safety Layer 5 — tracked count > 0
|
||||
if [[ "$TRACKED_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 tracked files — aborting to prevent mass deletion"
|
||||
notify "Radarr cleanup aborted on $(hostname) — 0 tracked files returned" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "$MOVIE_COUNT movies | $TRACKED_COUNT tracked movie files"
|
||||
|
||||
# Safety Layer 6 — percentage drop vs last known count, with rescan-aware retry.
|
||||
# ProcessMonitoredDownloads (this script's own pre-flight) is a different, lighter operation
|
||||
# than a full library rescan — but Radarr's own RescanMovie/DownloadedMoviesScan can be
|
||||
# triggered independently and would cause the exact same mid-scan count dip confirmed on
|
||||
# Lidarr 2026-07-16. Wait it out (calibrated to that command's own historical duration)
|
||||
# before treating a drop as genuine.
|
||||
_last_known=$(cat "$RADARR_TRACKED_COUNT_FILE" 2>/dev/null || echo 0)
|
||||
if [[ "$_last_known" -gt 0 ]]; then
|
||||
_strike=1
|
||||
while [[ "$_strike" -le 3 ]]; do
|
||||
_pct=$(awk "BEGIN {printf \"%d\", ($TRACKED_COUNT / $_last_known) * 100}")
|
||||
[[ "$_pct" -ge "${RADARR_MIN_TRACKED_PCT:-50}" ]] && break
|
||||
|
||||
_active_cmd=$(arr_active_rescan_command "radarr" "$RADARR_URL" "$RADARR_API_KEY" "v3")
|
||||
[[ -z "$_active_cmd" ]] && break # low count, nothing rescanning — genuine, don't retry
|
||||
|
||||
_wait=$(( $(arr_get_rescan_duration "radarr" "$_active_cmd" 300) / 2 ))
|
||||
[[ "$_wait" -lt 30 ]] && _wait=30
|
||||
warn "Tracked count ${_pct}% of last run, but $_active_cmd active — waiting ${_wait}s (strike ${_strike}/3)"
|
||||
sleep "$_wait"
|
||||
_fetch_tracked_files
|
||||
(( _strike++ ))
|
||||
done
|
||||
|
||||
if [[ "$_strike" -gt 3 ]]; then
|
||||
_active_cmd=$(arr_active_rescan_command "radarr" "$RADARR_URL" "$RADARR_API_KEY" "v3")
|
||||
if [[ -n "$_active_cmd" ]]; then
|
||||
warn "Radarr still busy ($_active_cmd) after 3 strikes — deferring to next scheduled run"
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
check_tracked_count_floor "$TRACKED_COUNT" "$RADARR_TRACKED_COUNT_FILE" "$RADARR_MIN_TRACKED_PCT" "Radarr Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Scan Movies Root ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Scanning Movies Root ━━━"
|
||||
info "Root: $RADARR_MOVIES_ROOT | Orphan age: ${RADARR_ORPHAN_AGE} days"
|
||||
|
||||
START=$(date +%s)
|
||||
ORPHAN_COUNT=0
|
||||
JUNK_COUNT=0
|
||||
RECENT_COUNT=0
|
||||
PROTECTED_COUNT=0
|
||||
ORPHAN_BYTES=0
|
||||
JUNK_BYTES=0
|
||||
|
||||
AGE_SECONDS=$(( RADARR_ORPHAN_AGE * 86400 ))
|
||||
NOW=$(date +%s)
|
||||
|
||||
# Files classified ORPHAN/JUNK below get their path recorded here, so the deletion pass can
|
||||
# just delete them directly instead of re-walking and re-classifying every SCAN_ROOTS entry a
|
||||
# second time (2026-07-17) — the size-threshold check below needs to know the total before
|
||||
# deleting anything, not before knowing what to delete.
|
||||
# Carries size and ctime alongside the path now, because the budget pass below has to order by
|
||||
# age and stop at a byte ceiling — neither of which a bare path list can answer.
|
||||
TO_DELETE_FILE="$TMP_DIR/to_delete_paths.txt"
|
||||
> "$TO_DELETE_FILE"
|
||||
|
||||
# ── Orphan strikes ────────────────────────────────────────────────────────────────────────────
|
||||
# A file must classify for deletion on RADARR_ORPHAN_STRIKE_LIMIT consecutive runs before it is
|
||||
# actually removed. Gate 6 already refuses a run whose tracked count collapsed; this covers the
|
||||
# partial failure underneath that threshold — one root folder failing to enumerate makes its
|
||||
# movies look orphaned while the overall percentage still looks fine, and a transient fault will
|
||||
# not reproduce on the next run.
|
||||
#
|
||||
# The file is REBUILT from this run's classifications rather than edited in place, which is what
|
||||
# prunes it: anything that stopped being an orphan simply is not written again, so a file that
|
||||
# Radarr re-adopts loses its strikes without needing a reset pass to find it.
|
||||
#
|
||||
# Keyed by host path, which is why this could not have worked before 2026-08-26 — wd_state_set
|
||||
# built a regex from the key, and a release tag like [Bluray-1080p] holds the reversed range 1-0,
|
||||
# so every write truncated the store to one line. See common.sh.
|
||||
RADARR_ORPHAN_STRIKE_LIMIT="${RADARR_ORPHAN_STRIKE_LIMIT:-2}"
|
||||
STRIKES_FILE="${RADARR_ORPHAN_STRIKES_FILE:-$DB_DIR/radarr_orphan_strikes.tsv}"
|
||||
mkdir -p "$(dirname "$STRIKES_FILE")" 2>/dev/null || true
|
||||
touch "$STRIKES_FILE" 2>/dev/null || true
|
||||
STRIKES_NEW="$TMP_DIR/strikes_new.tsv"
|
||||
> "$STRIKES_NEW"
|
||||
HELD_COUNT=0
|
||||
HELD_BYTES=0
|
||||
|
||||
# Records this run's strike for a file and says whether it has served enough of them.
|
||||
# Returns 0 when the file may be deleted, 1 when it is still accruing.
|
||||
orphan_strike_ok() {
|
||||
local path="$1" prev strikes
|
||||
prev=$(wd_state_get "$path" "$STRIKES_FILE"); prev="${prev//[^0-9]/}"
|
||||
strikes=$(( ${prev:-0} + 1 ))
|
||||
printf '%s:%s\n' "$path" "$strikes" >> "$STRIKES_NEW"
|
||||
(( strikes >= RADARR_ORPHAN_STRIKE_LIMIT )) && return 0
|
||||
warn " strike $strikes/$RADARR_ORPHAN_STRIKE_LIMIT — not removing yet: $path"
|
||||
return 1
|
||||
}
|
||||
|
||||
while read -r FILE_SIZE FILE_CTIME filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
FILE_CTIME="${FILE_CTIME%%.*}"
|
||||
|
||||
if [[ -n "${TRACKED_MAP[$filepath]:-}" ]]; then
|
||||
log "TRACKED: $filepath"
|
||||
continue
|
||||
fi
|
||||
|
||||
if matches_pattern_list "$filepath" "${RADARR_PROTECTED_PATTERNS[@]}"; then
|
||||
log "$ICON_PROTECTED PROTECTED: $filepath"
|
||||
(( PROTECTED_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
if has_extension "$filepath" "${RADARR_EXTENSIONS[@]}"; then
|
||||
# ctime, not mtime — an import preserves the release's original mtime, so a file
|
||||
# Radarr moved in today can read as years old and skip this gate entirely.
|
||||
# Measured 2026-07-27: 400 of 400 files imported that week had mtimes over 7
|
||||
# days, one of them 9613 days. ctime is stamped when the file lands on this
|
||||
# filesystem and cannot be carried in from an archive. This only holds because
|
||||
# media_shares_permissions.sh applies owner/mode conditionally — a blanket
|
||||
# chown/chmod restamps every inode nightly and would peg every file at age 0.
|
||||
FILE_AGE=$(( NOW - FILE_CTIME ))
|
||||
|
||||
if [[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && [[ "$SKIP_AGE_CHECK" != true ]]; then
|
||||
log "RECENT (skipping): $filepath"
|
||||
(( RECENT_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
warn "$ICON_TRASH ORPHAN: $filepath"
|
||||
(( ORPHAN_COUNT++ ))
|
||||
ORPHAN_BYTES=$(( ORPHAN_BYTES + FILE_SIZE ))
|
||||
if ! orphan_strike_ok "$filepath"; then (( HELD_COUNT++ )); HELD_BYTES=$(( HELD_BYTES + FILE_SIZE )); continue; fi
|
||||
printf '%s\t%s\t%s\n' "$FILE_SIZE" "$FILE_CTIME" "$filepath" >> "$TO_DELETE_FILE"
|
||||
else
|
||||
log "JUNK: $filepath"
|
||||
(( JUNK_COUNT++ ))
|
||||
JUNK_BYTES=$(( JUNK_BYTES + FILE_SIZE ))
|
||||
if ! orphan_strike_ok "$filepath"; then (( HELD_COUNT++ )); HELD_BYTES=$(( HELD_BYTES + FILE_SIZE )); continue; fi
|
||||
printf '%s\t%s\t%s\n' "$FILE_SIZE" "$FILE_CTIME" "$filepath" >> "$TO_DELETE_FILE"
|
||||
fi
|
||||
|
||||
# -printf gets size + mtime directly from find's own stat() during the walk, instead of a
|
||||
# separate stat fork per file (2026-07-17) — measured ~130x faster per file (0.033ms vs
|
||||
# 4.3ms), since find already has to stat() every entry anyway to know it's -type f.
|
||||
done < <(
|
||||
for host_path in "${SCAN_ROOTS[@]}"; do
|
||||
[[ -d "$host_path" ]] && find "$host_path" -type f -printf '%s %C@ %p\n' 2>/dev/null
|
||||
done | sort -u
|
||||
)
|
||||
|
||||
# Eligible, not classified. A file still serving its strikes was counted as an orphan above — it
|
||||
# is one — but it is not going to be deleted this run, so it must not appear in the denominator
|
||||
# the budget reports against or the run claims to have skipped work it never queued.
|
||||
TOTAL_DELETE_BYTES=$(( ORPHAN_BYTES + JUNK_BYTES - HELD_BYTES ))
|
||||
TOTAL_REMOVED=$(( ORPHAN_COUNT + JUNK_COUNT - HELD_COUNT ))
|
||||
|
||||
# Rebuilt, never edited: a path absent from this run is absent from the file, so a file Radarr
|
||||
# re-adopts drops its strikes with no reset pass needed. Skipped on a dry run — a preview that
|
||||
# advanced real strike counters would make the next real run delete a run early.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
mv "$STRIKES_NEW" "$STRIKES_FILE" 2>/dev/null || warn "Could not update $STRIKES_FILE"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 7 — Deletion Size Threshold ━━━
|
||||
# ==============================================================================================
|
||||
# The ceiling is a per-run budget, not a veto. It still means what it always meant — no single run
|
||||
# removes more than RADARR_MAX_DELETE_GB — but a backlog larger than the ceiling now drains over
|
||||
# consecutive nights instead of failing the orchestrator forever on a queue it cannot clear.
|
||||
# ── AI note (AI_ASSIST_CLEANUP) ───────────────────────────────────────────────────────────────
|
||||
# Describes the shape of what was classified. It decides nothing: the eligible set, the budget and
|
||||
# the strikes are all settled above and none of them read this. Switch AI_ASSIST_CLEANUP off and
|
||||
# the run removes exactly the same files — the log just loses a paragraph.
|
||||
#
|
||||
# ctime clustering is the signal worth surfacing. A normal upgrade cycle dribbles in over weeks; a
|
||||
# lump sharing one narrow ctime window with mtimes spread across months is a bulk write-back, which
|
||||
# is what a partnership merge against a partner holding older copies produces. That distinction
|
||||
# took a person an evening on 2026-08-26 and is the whole reason this note exists.
|
||||
if [[ "$ORPHAN_COUNT" -gt 0 ]] && [[ -s "$TO_DELETE_FILE" ]]; then
|
||||
_ai_ev=$(awk -F'\t' '
|
||||
{ n++; bytes += $1
|
||||
c = int($2)
|
||||
if (cmin == 0 || c < cmin) cmin = c
|
||||
if (c > cmax) cmax = c
|
||||
bucket[int(c / 21600)]++ }
|
||||
END {
|
||||
for (b in bucket) if (bucket[b] > top) { top = bucket[b] }
|
||||
printf "files=%d bytes_gb=%.1f ctime_span_hours=%.1f largest_6h_ctime_bucket=%d\n",
|
||||
n, bytes/1073741824, (cmax-cmin)/3600, top
|
||||
}' "$TO_DELETE_FILE")
|
||||
_ai_mt=$(cut -d"$(printf '\t')" -f3 "$TO_DELETE_FILE" | head -8 \
|
||||
| while IFS= read -r p; do [[ -f "$p" ]] && \
|
||||
printf '%s %s\n' "$(stat -c %y "$p" 2>/dev/null | cut -c1-7)" "$(basename "$p")"; done)
|
||||
|
||||
_ai_note=$(ai_assist_note AI_ASSIST_CLEANUP "You are looking at files an automated media-library cleanup has classified for deletion on an Unraid server. They are files on disk that the Radarr database no longer references.
|
||||
|
||||
EVIDENCE
|
||||
$_ai_ev
|
||||
sample (modification month, then path):
|
||||
$_ai_mt
|
||||
|
||||
A normal quality-upgrade cycle produces orphans whose ctimes are spread out over weeks, because each upgrade happens on its own day. A bulk event - a sync or restore writing files back onto this host - produces orphans sharing one narrow ctime window while their modification times stay spread across months, because the copy preserves modification time but resets ctime.
|
||||
|
||||
In no more than three sentences, say which of those two this looks like and name the numbers above that support it. Do not recommend an action. Do not speculate beyond the evidence given.") || _ai_note=""
|
||||
|
||||
if [[ -n "$_ai_note" ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR AI note on this classification ━━━"
|
||||
printf '%s\n' "$_ai_note"
|
||||
fi
|
||||
unset _ai_ev _ai_mt
|
||||
fi
|
||||
|
||||
BUDGET_FILE="$TMP_DIR/to_delete_budgeted.txt"
|
||||
|
||||
if [[ "$I_KNOW" == true ]]; then
|
||||
warn "OVERRIDE — --i-know-what-im-doing active, per-run budget not applied"
|
||||
cut -d"$(printf '\t')" -f3- "$TO_DELETE_FILE" > "$BUDGET_FILE"
|
||||
_BUDGET_KEPT_COUNT=$TOTAL_REMOVED; _BUDGET_KEPT_BYTES=$TOTAL_DELETE_BYTES
|
||||
_BUDGET_DEFERRED_COUNT=0; _BUDGET_DEFERRED_BYTES=0; _BUDGET_STUCK=""
|
||||
else
|
||||
apply_delete_budget "$TO_DELETE_FILE" "$BUDGET_FILE" "$RADARR_MAX_DELETE_GB"
|
||||
|
||||
if [[ -n "$_BUDGET_STUCK" ]]; then
|
||||
# One file larger than the whole budget can never fit, so it would be re-found and
|
||||
# re-deferred every night. Name it rather than loop on it silently.
|
||||
error "Single file exceeds the ${RADARR_MAX_DELETE_GB}GB budget on its own — nothing removed this run"
|
||||
error " $_BUDGET_STUCK"
|
||||
error "Raise RADARR_MAX_DELETE_GB or clear this one with --i-know-what-im-doing"
|
||||
notify "Radarr cleanup stalled on $(hostname) — one file exceeds the ${RADARR_MAX_DELETE_GB}GB budget" \
|
||||
"Radarr Cleanup" "warning"
|
||||
elif [[ "$_BUDGET_DEFERRED_COUNT" -gt 0 ]]; then
|
||||
warn "Budget ${RADARR_MAX_DELETE_GB}GB — removing $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED ($(format_bytes "$_BUDGET_KEPT_BYTES")), deferring $_BUDGET_DEFERRED_COUNT ($(format_bytes "$_BUDGET_DEFERRED_BYTES")) to the next run"
|
||||
warn "Oldest first — the deferred files are the newest and are re-evaluated tomorrow"
|
||||
notify "Radarr cleanup removed $(format_bytes "$_BUDGET_KEPT_BYTES") of $(format_bytes "$TOTAL_DELETE_BYTES") on $(hostname) — $_BUDGET_DEFERRED_COUNT file(s) deferred to the next run" \
|
||||
"Radarr Cleanup" "normal"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Execute Deletions ─────────────────────────────────────────────────────────────────────────
|
||||
# Reuses TO_DELETE_FILE from the classification pass above instead of re-walking and
|
||||
# re-classifying every SCAN_ROOTS entry again.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
rm -f "$filepath" 2>/dev/null || error "Failed to delete: $filepath"
|
||||
done < "$BUDGET_FILE"
|
||||
|
||||
info "Cleaning up empty folders..."
|
||||
for host_path in "${SCAN_ROOTS[@]}"; do
|
||||
[[ -d "$host_path" ]] && \
|
||||
find "$host_path" -mindepth 1 -type d -empty -delete 2>/dev/null
|
||||
done
|
||||
info "Empty folders removed"
|
||||
fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
ORPHAN_HUMAN=$(format_bytes "$ORPHAN_BYTES")
|
||||
JUNK_HUMAN=$(format_bytes "$JUNK_BYTES")
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY RADARR CLEANUP SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Tracked: $TRACKED_COUNT files ($MOVIE_COUNT movies)"
|
||||
echo "$ICON_SHIELD Protected: $PROTECTED_COUNT files (artwork, subtitles, metadata)"
|
||||
echo "$ICON_TRASH Orphans: $ORPHAN_COUNT files ($ORPHAN_HUMAN)"
|
||||
echo "$ICON_TRASH Junk: $JUNK_COUNT files ($JUNK_HUMAN)"
|
||||
echo "$ICON_SKIP Recent skipped: $RECENT_COUNT files (under ${RADARR_ORPHAN_AGE} days)"
|
||||
[[ "${HELD_COUNT:-0}" -gt 0 ]] && \
|
||||
echo "$ICON_SKIP Held (strikes): $HELD_COUNT files ($(format_bytes "$HELD_BYTES")) — under ${RADARR_ORPHAN_STRIKE_LIMIT} consecutive runs"
|
||||
[[ "${_BUDGET_DEFERRED_COUNT:-0}" -gt 0 ]] && \
|
||||
echo "$ICON_SKIP Deferred: $_BUDGET_DEFERRED_COUNT files ($(format_bytes "$_BUDGET_DEFERRED_BYTES")) — over the ${RADARR_MAX_DELETE_GB}GB run budget"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no files deleted"
|
||||
elif [[ "$TOTAL_REMOVED" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Clean — nothing to remove"
|
||||
else
|
||||
# What was actually removed, not what was classified. With a budget in force those differ, and
|
||||
# reporting the classification as the outcome is the oldest bug shape in this codebase.
|
||||
warn "$ICON_DONE Removed $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED classified files ($(format_bytes "$_BUDGET_KEPT_BYTES"))"
|
||||
notify "Radarr cleanup on $(hostname) — removed $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED classified files ($(format_bytes "$_BUDGET_KEPT_BYTES"))$([[ "${_BUDGET_DEFERRED_COUNT:-0}" -gt 0 ]] && echo ", $_BUDGET_DEFERRED_COUNT deferred")" \
|
||||
"Radarr Cleanup" "warning"
|
||||
# Notify Emby to clean missing files — removes ghost entries immediately
|
||||
notify_emby_scan
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# Write stats for sunday_morning_coffee_report.sh
|
||||
if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_CLEANUP_STATS:-}" ]]; then
|
||||
echo "$(date '+%Y-%m-%d')|radarr|${ORPHAN_COUNT}|${ORPHAN_BYTES}|${JUNK_COUNT}|${JUNK_BYTES}|${RECENT_COUNT}|${TRACKED_COUNT}" \
|
||||
>> "$ARR_CLEANUP_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
Regular → Executable
+31
-5
@@ -10,6 +10,10 @@
|
||||
# or downloaded. Most are announced-but-never-released films delisted before
|
||||
# release.
|
||||
#
|
||||
# Cache-first movie list (2026-07-17) — comes from the shared tracked-data cache via
|
||||
# arr_get_tracked_data(), fresh (kept warm every 30min by arr_cache_prefill.sh), live fetch
|
||||
# as fallback.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
@@ -24,6 +28,25 @@
|
||||
# Per-deletion output is always visible — deletions are never silently swallowed.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Health Error Hygiene
|
||||
# status="deleted" entries can never be monitored or downloaded — they only
|
||||
# generate persistent health errors. Removing them is maintenance, not
|
||||
# data loss: the content never existed on disk for most of these entries.
|
||||
#
|
||||
# Conservative File Handling
|
||||
# Files are not deleted by default because most TMDb-removed entries are
|
||||
# announced-but-never-released films with no files. The --delete-files flag
|
||||
# is an explicit opt-in, not the default path.
|
||||
#
|
||||
# Exclusion List Prevents Re-add
|
||||
# Removed entries are added to Radarr's import exclusion list by default.
|
||||
# Without this, the same deleted entry can be re-added by lists or searches
|
||||
# and immediately generate the same health error again.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -95,12 +118,13 @@ acquire_lock
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "$RADARR_URL" ]]; then
|
||||
log "Radarr not configured for $MY_ID — nothing to do"
|
||||
echo "Radarr not configured for $MY_ID — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ADD_EXCLUSION="${RADARR_DROPPED_ADD_EXCLUSION:-true}"
|
||||
|
||||
log "$ICON_GEAR Config: url=${RADARR_URL} add-exclusion=${ADD_EXCLUSION} delete-files=${DELETE_FILES:-false}"
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $RADARR_URL"
|
||||
[[ "$DELETE_FILES" == true ]] && warn "DELETE FILES MODE — files will be removed from disk"
|
||||
[[ "$DELETE_FILES" == false ]] && echo " Files: records only (use --delete-files to also remove from disk)"
|
||||
@@ -133,8 +157,10 @@ if ! curl -sf --connect-timeout 5 --max-time 10 \
|
||||
exit 1
|
||||
fi
|
||||
|
||||
MOVIES=$(curl -sf --connect-timeout 5 --max-time 30 \
|
||||
"$RADARR_URL/api/v3/movie?apikey=$RADARR_API_KEY" 2>/dev/null)
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (kept current
|
||||
# every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to a live
|
||||
# fetch when it's stale, and waits out an active rescan before either.
|
||||
MOVIES=$(arr_get_tracked_data "radarr" "$RADARR_URL" "$RADARR_API_KEY" "v3")
|
||||
|
||||
if [[ -z "$MOVIES" || "$MOVIES" == "null" ]]; then
|
||||
error "Radarr movie API returned empty"
|
||||
@@ -146,7 +172,7 @@ DROPPED=$(echo "$MOVIES" | jq '[.[] | select(.status == "deleted")] | length')
|
||||
echo " $TOTAL movies total — $DROPPED dropped from TMDb"
|
||||
|
||||
if [[ "$DROPPED" -eq 0 ]]; then
|
||||
log "No TMDb-removed movies found — nothing to do"
|
||||
echo "No TMDb-removed movies found — nothing to do"
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY RADARR TMDB REMOVED SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
@@ -197,7 +223,7 @@ while IFS=$'\t' read -r id title year tmdb_id has_file file_size; do
|
||||
CURL_EXIT=$?
|
||||
|
||||
if [[ "$CURL_EXIT" -eq 0 ]]; then
|
||||
log " Removed from Radarr ✅"
|
||||
echo " Removed from Radarr ✅"
|
||||
REMOVED+=("$title")
|
||||
if [[ "$DELETE_PARAM" == "true" ]]; then
|
||||
(( FILES_DELETED++ ))
|
||||
Executable
+557
@@ -0,0 +1,557 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================ Sonarr Content Classification Scan ==============================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Same problem as radarr_classification_scan.sh, TV side: Overseerr lets any user request a
|
||||
# show into the wrong root folder (kids shows added to the general TV share, anime added to
|
||||
# Kids_Tv_Shows, etc.). This script reads Sonarr's tracked series list and classifies every
|
||||
# series as anime / kids-only / regular using metadata signals alone (genre, certification,
|
||||
# network, original language) — then reports where a series' computed classification
|
||||
# disagrees with the root folder it's actually sitting in, in both directions:
|
||||
#
|
||||
# FORWARD — a series classified as anime/kids is sitting outside its dedicated root
|
||||
# REVERSE — a series sitting inside the kids/anime root doesn't match that classification
|
||||
#
|
||||
# Report-only by default. Every rule below was validated against this library's real data
|
||||
# before being adopted — see the companion comment block in master.conf above the curated
|
||||
# lists. Pass --move to actually act (see MOVE MODE below) — nothing writes to Sonarr unless
|
||||
# that flag is given.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CLASSIFICATION RULES — DIFFERENT FIELD MODEL THAN RADARR, NOT A COPY-PASTE
|
||||
# ==============================================================================================
|
||||
#
|
||||
# TV metadata (TheTVDB, via Sonarr) shapes these signals differently than movie metadata
|
||||
# (TMDb, via Radarr) — every difference below was confirmed live, not assumed:
|
||||
# - Sonarr has an explicit "Anime" genre tag; Radarr does not.
|
||||
# - Sonarr uses a single "network" field (TheTVDB's broadcaster), not "studio".
|
||||
# - Certification is on the US TV Parental Guidelines scale (TV-Y/TV-Y7/TV-G/TV-PG/
|
||||
# TV-14/TV-MA), not the MPAA scale — the tiers do not mean the same thing at the same
|
||||
# position (TV-G is "general audience", not "for children", unlike movie G).
|
||||
#
|
||||
# is_anime:
|
||||
# genre "Anime" (corroborated by Japanese language OR a Japan network — the bare tag alone
|
||||
# produced a real false positive: "Craig of the Creek", an all-American Cartoon Network
|
||||
# show, carries an "Anime" genre tag on TheTVDB for no discernible reason)
|
||||
# OR (genre Animation AND originalLanguage Japanese)
|
||||
# OR network in SONARR_ANIME_NETWORKS
|
||||
# Always wins over kids when both could apply — explicit priority, not a tiebreak.
|
||||
#
|
||||
# is_kids ("kids will end up watching this alone" — NOT "family show night"):
|
||||
# not is_anime AND (
|
||||
# genre "Children" (NOT "Family" — see below)
|
||||
# OR certification in (TV-Y, TV-Y7) (NOT TV-G — see below)
|
||||
# OR network in SONARR_KIDS_NETWORKS
|
||||
# )
|
||||
# "Family" genre and "TV-G" certification were both tested standalone and rejected —
|
||||
# both catch general-audience live-action content the whole household watches together
|
||||
# (I Love Lucy, The Brady Bunch, Full House, Homestead Rescue), not kids-only content.
|
||||
# Blanket "Animation" genre was also tested and rejected — it's dominated on TV by adult
|
||||
# animated sitcoms (Rick and Morty, BoJack Horseman, Family Guy, South Park), unlike the
|
||||
# movie side where it's a usable (gated) signal.
|
||||
#
|
||||
# No junk-detection tier here (unlike Radarr) — TheTVDB's ratings/imdbId data is far
|
||||
# sparser than TMDb's even for completely legitimate shows (confirmed live: "The Pussycat
|
||||
# Dolls Present: The Search for the Next Doll", a real 2007 MTV show, has ratings.votes=0
|
||||
# and imdbId=null) — the vote-count heuristic that works for Radarr would flag real content
|
||||
# for removal here, so it's deliberately not reused. --remove-junk from the Radarr script has
|
||||
# no Sonarr equivalent for the same reason.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Report pass (always):
|
||||
# check_api → check_arr_version → arr_get_tracked_data (cache-first, one call)
|
||||
# → classify every series → report FORWARD and REVERSE disagreements → exit
|
||||
#
|
||||
# Move pass (--move only), described in detail below:
|
||||
#
|
||||
# Acts on FORWARD misplacements (classified anime/kids, sitting in the wrong root) and on
|
||||
# REVERSE-KIDS leaks (adult certification sitting in the kids root — moved back to
|
||||
# SONARR_GENERAL_ROOT). Does NOT act on REVERSE-ANIME leaks — those are genuine judgment
|
||||
# calls, since deliberate style placements (Castlevania-type Western/Chinese animation
|
||||
# grouped with anime by choice) legitimately live in the anime root without matching the
|
||||
# anime signal.
|
||||
#
|
||||
# episodeFileCount is Sonarr's equivalent of Radarr's hasFile — a series can have 0 files
|
||||
# (fully monitored, nothing downloaded) even while correctly classified. Those get their
|
||||
# rootFolderPath/path corrected and an immediate SeriesSearch triggered rather than a file
|
||||
# move (mirrors radarr_classification_scan.sh's handling of hasFile=false movies).
|
||||
#
|
||||
# One series at a time, verified after each. moveFiles=true flips the DB (rootFolderPath/
|
||||
# episodeFileCount) instantly, but the physical move is a separate async MoveSeries command
|
||||
# Sonarr drains one at a time internally — DB fields alone can report "moved" while the real
|
||||
# files are still sitting at the old path behind other queued moves (confirmed live: "Full
|
||||
# House" reported episodeFileCount:192 at the new path via API while the actual 75GB/192
|
||||
# files hadn't moved yet). Each move polls its own MoveSeries command to "completed" before
|
||||
# the DB-field check runs, so a batch can't compound the race the way a bare sleep-and-check did.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Curated Lists, Not Bare Genre/Cert Matching — see master.conf comments for exclusions.
|
||||
# Cache-First — arr_get_tracked_data() same as sonarr_cleanup.sh, single call regardless
|
||||
# of library size. Refreshed after --move writes so no other script reads stale data.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Required by the container interaction and state writes.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents a scheduled run overlapping a manual --move.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases SONARR_URL / SONARR_API_KEY / the root literals.
|
||||
#
|
||||
# curl + jq Dependency Check
|
||||
# Fails fast if either is missing — every classification signal is parsed with jq,
|
||||
# and a missing jq would evaluate each signal to empty and classify nothing.
|
||||
#
|
||||
# Report-Only Default
|
||||
# Nothing is written to Sonarr without --move. The scan is safe to schedule and
|
||||
# safe to run repeatedly while tuning the curated lists.
|
||||
#
|
||||
# Required Var Check
|
||||
# require_var on SONARR_URL and SONARR_API_KEY before any request.
|
||||
#
|
||||
# API Reachability + Version Gate
|
||||
# check_api then check_arr_version against SONARR_VERSION_MAJOR. A major version
|
||||
# bump can move or rename the fields every rule here depends on, so a mismatch
|
||||
# aborts rather than classifying against an unknown schema.
|
||||
#
|
||||
# Empty Library Abort
|
||||
# A response of 0 series aborts. An empty list is indistinguishable from "nothing
|
||||
# is misplaced" and would otherwise report a clean library during an API fault.
|
||||
#
|
||||
# Unconfigured Root Skip
|
||||
# A blank SONARR_GENERAL_ROOT / KIDS_ROOT / ANIME_ROOT skips that category's
|
||||
# checks rather than erroring — a host with no dedicated root is a valid setup,
|
||||
# and a blank value must never be compared against as if it were a real path.
|
||||
#
|
||||
# One At A Time, Stop On First Failure
|
||||
# Series are moved individually and the batch halts on the first failure rather
|
||||
# than continuing. A misclassified root or a failing move is a condition to
|
||||
# review, not to repeat across the library.
|
||||
#
|
||||
# Async Move Completion Polling
|
||||
# moveFiles=true flips the DB instantly while the physical move is a separate
|
||||
# async MoveSeries command Sonarr drains one at a time. Each move locates its own
|
||||
# command and polls it to "completed" (bounded by SONARR_MOVE_POLL_TIMEOUT) before
|
||||
# anything else is checked. Without this a batch reports every series moved while
|
||||
# the files are still queued at the old path — and downstream orphan cleanup can
|
||||
# act on that gap.
|
||||
#
|
||||
# Post-Move Re-Verification
|
||||
# The PUT response is never trusted. The series is re-fetched and both
|
||||
# rootFolderPath and episodeFileCount are confirmed against expectations before
|
||||
# the move counts as successful.
|
||||
#
|
||||
# Reverse-Anime Leaks Excluded From Moves
|
||||
# Deliberate style placements (Western/Chinese animation grouped with anime by
|
||||
# choice) legitimately sit in the anime root. Those are reported, never moved.
|
||||
#
|
||||
# Post-Write Cache Refresh
|
||||
# The tracked-data cache is refreshed after --move writes so no other arr script
|
||||
# reads a stale rootFolderPath.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
# SONARR_URL / SONARR_API_KEY / SONARR_TV_ROOT — existing, aliased by detect_hosts()
|
||||
# SONARR_GENERAL_ROOT / SONARR_KIDS_ROOT / SONARR_ANIME_ROOT — rootFolderPath literals as
|
||||
# reported by the API (e.g. "/tv", "/kids tv", "/ext-anime-shows") — leave blank on a
|
||||
# host with no dedicated root for that category; the corresponding checks are skipped,
|
||||
# not treated as an error.
|
||||
#
|
||||
# master.conf
|
||||
# SONARR_ANIME_NETWORKS / SONARR_KIDS_NETWORKS — curated network allowlists
|
||||
# SONARR_VERSION_MAJOR — expected API major version (reused from sonarr_cleanup.sh)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# sonarr_classification_scan.sh — normal run, prints report
|
||||
# sonarr_classification_scan.sh --log — verbose (per-series list)
|
||||
# sonarr_classification_scan.sh --status — show config and exit
|
||||
# sonarr_classification_scan.sh --move — act on forward misplacements + reverse-kids-leak
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# --move is a script-local flag, not one parse_args recognizes — check the raw args before
|
||||
# they get filtered into PARSED_ARGS.
|
||||
MOVE_MODE=false
|
||||
for _arg in "$@"; do
|
||||
[[ "$_arg" == "--move" ]] && MOVE_MODE=true
|
||||
done
|
||||
unset _arg
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "${SONARR_URL:-}" ]] || [[ -z "${SONARR_API_KEY:-}" ]]; then
|
||||
info "Sonarr not configured on $MY_ID ($LOCAL_SERVER_NAME) — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
require_var SONARR_URL
|
||||
require_var SONARR_API_KEY
|
||||
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Sonarr URL: $SONARR_URL"
|
||||
echo "$ICON_GEAR TV root: $SONARR_TV_ROOT"
|
||||
echo "$ICON_GEAR General root: ${SONARR_GENERAL_ROOT:-<not configured>}"
|
||||
echo "$ICON_GEAR Kids root: ${SONARR_KIDS_ROOT:-<not configured>}"
|
||||
echo "$ICON_GEAR Anime root: ${SONARR_ANIME_ROOT:-<not configured>}"
|
||||
echo "$ICON_GEAR Anime networks: ${#SONARR_ANIME_NETWORKS[@]} curated"
|
||||
echo "$ICON_GEAR Kids networks: ${#SONARR_KIDS_NETWORKS[@]} curated"
|
||||
echo "$ICON_GEAR Move mode: $MOVE_MODE"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching Sonarr Library ━━━"
|
||||
|
||||
if ! check_api "$SONARR_URL" "Sonarr" 10; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
check_arr_version "$SONARR_URL" "$SONARR_API_KEY" "v3" "$SONARR_VERSION_MAJOR" "Sonarr" || exit 1
|
||||
|
||||
SERIES_RESPONSE=$(arr_get_tracked_data "sonarr" "$SONARR_URL" "$SONARR_API_KEY" "v3") || {
|
||||
error "Failed to fetch series from Sonarr"
|
||||
exit 1
|
||||
}
|
||||
|
||||
SERIES_COUNT=$(echo "$SERIES_RESPONSE" | jq -r 'length' 2>/dev/null)
|
||||
if [[ -z "$SERIES_COUNT" ]] || [[ "$SERIES_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 series — aborting"
|
||||
exit 1
|
||||
fi
|
||||
info "$SERIES_COUNT series loaded"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Classify ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Classifying ━━━"
|
||||
|
||||
ANIME_NETWORKS_JSON=$(printf '%s\n' "${SONARR_ANIME_NETWORKS[@]}" | jq -R . | jq -s .)
|
||||
KIDS_NETWORKS_JSON=$(printf '%s\n' "${SONARR_KIDS_NETWORKS[@]}" | jq -R . | jq -s .)
|
||||
# Not just the US TV-MA/TV-14 tiers — non-US certification scales use different labels for the
|
||||
# same "clearly adult" tier (confirmed live: "Tomb Raider: The Legend of Lara Croft" is "16+").
|
||||
ADULT_CERT_JSON='["TV-MA","TV-14","MA15+","16","16+","18","15","14"]'
|
||||
|
||||
RESULTS=$(echo "$SERIES_RESPONSE" | jq \
|
||||
--argjson animeNetworks "$ANIME_NETWORKS_JSON" \
|
||||
--argjson kidsNetworks "$KIDS_NETWORKS_JSON" \
|
||||
--argjson adultCert "$ADULT_CERT_JSON" \
|
||||
--arg animeRoot "${SONARR_ANIME_ROOT:-}" \
|
||||
--arg kidsRoot "${SONARR_KIDS_ROOT:-}" '
|
||||
def is_anime:
|
||||
(any(.genres[]?; . == "Anime")
|
||||
and (.originalLanguage.name == "Japanese" or (.network as $n | $animeNetworks | index($n) != null)))
|
||||
or (any(.genres[]?; . == "Animation") and .originalLanguage.name == "Japanese")
|
||||
or (.network as $n | $animeNetworks | index($n) != null);
|
||||
def is_kids:
|
||||
(is_anime | not) and (
|
||||
any(.genres[]?; . == "Children")
|
||||
or (.certification as $c | ["TV-Y","TV-Y7"] | index($c) != null)
|
||||
or (.network as $n | $kidsNetworks | index($n) != null)
|
||||
);
|
||||
|
||||
map(
|
||||
{
|
||||
title, id, network, certification, rootFolderPath,
|
||||
episodeFileCount: (.statistics.episodeFileCount // 0),
|
||||
is_anime: is_anime,
|
||||
is_kids: is_kids
|
||||
} |
|
||||
. + {
|
||||
forward_anime_miss: (.is_anime and $animeRoot != "" and .rootFolderPath != $animeRoot),
|
||||
forward_kids_miss: (.is_kids and $kidsRoot != "" and .rootFolderPath != $kidsRoot),
|
||||
reverse_anime_leak: ((.is_anime | not) and $animeRoot != "" and .rootFolderPath == $animeRoot),
|
||||
reverse_kids_leak: ((.is_anime | not) and (.is_kids | not) and $kidsRoot != "" and .rootFolderPath == $kidsRoot
|
||||
and (.certification as $c | $adultCert | index($c) != null))
|
||||
}
|
||||
)
|
||||
')
|
||||
|
||||
FORWARD_ANIME_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.forward_anime_miss)] | length')
|
||||
FORWARD_KIDS_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.forward_kids_miss)] | length')
|
||||
REVERSE_ANIME_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.reverse_anime_leak)] | length')
|
||||
REVERSE_KIDS_COUNT=$(echo "$RESULTS" | jq '[.[] | select(.reverse_kids_leak)] | length')
|
||||
|
||||
if [[ "$ENABLE_LOGGING" == true ]]; then
|
||||
echo "$RESULTS" | jq -r '.[] | select(.forward_anime_miss or .forward_kids_miss or .reverse_anime_leak or .reverse_kids_leak) |
|
||||
" [\(if .forward_anime_miss then "FORWARD-ANIME" elif .forward_kids_miss then "FORWARD-KIDS" elif .reverse_anime_leak then "REVERSE-ANIME" elif .reverse_kids_leak then "REVERSE-KIDS" else "?" end)] \(.title) (root: \(.rootFolderPath), network: \(.network // "n/a"), cert: \(.certification // "n/a"))"'
|
||||
fi
|
||||
|
||||
# ── The reverse leaks, written down ───────────────────────────────────────────────────────────
|
||||
# This block's own header says it does not act on REVERSE-ANIME leaks because they are genuine
|
||||
# judgement calls. That is right, and it is also why they are the one result worth persisting:
|
||||
# every other bucket either self-resolves or is acted on by --move, while these accumulate as a
|
||||
# number in a summary nobody can do anything with. Seventeen of them hid two live-action dramas
|
||||
# filed under anime for as long as the count stayed a count.
|
||||
#
|
||||
# Written as the script's own verdict so anything reading it — the triage that reads this next —
|
||||
# inherits the classification rather than computing a second opinion from the same metadata.
|
||||
# Report-only: this records what was found, it does not change what happens to any of it.
|
||||
if [[ -n "${STATE_DIR:-}" ]] && [[ "$DRY_RUN" != true ]]; then
|
||||
_review_file="$STATE_DIR/arr_classification_review.json"
|
||||
echo "$RESULTS" | jq -c --arg host "$MY_ID" --argjson ts "$(date +%s)" '
|
||||
{ host: $host, ts: $ts, arr: "sonarr",
|
||||
reverse_anime: [ .[] | select(.reverse_anime_leak) |
|
||||
{ title, root: .rootFolderPath, network: (.network // ""), cert: (.certification // ""),
|
||||
lang: (.originalLanguage.name // .originalLanguage // ""), id: .id } ],
|
||||
reverse_kids: [ .[] | select(.reverse_kids_leak) |
|
||||
{ title, root: .rootFolderPath, network: (.network // ""), cert: (.certification // "") } ] }
|
||||
' > "$_review_file" 2>/dev/null \
|
||||
&& log "$ICON_GEAR Review list written — $REVERSE_ANIME_COUNT anime leak(s) for triage" \
|
||||
|| warn "Could not write $_review_file — triage will have nothing to read"
|
||||
unset _review_file
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY SONARR CLASSIFICATION SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Series scanned: $SERIES_COUNT"
|
||||
echo "$ICON_TRASH Forward — anime miss: $FORWARD_ANIME_COUNT (classified anime, outside ${SONARR_ANIME_ROOT:-<unconfigured>})"
|
||||
echo "$ICON_TRASH Forward — kids miss: $FORWARD_KIDS_COUNT (classified kids, outside ${SONARR_KIDS_ROOT:-<unconfigured>})"
|
||||
echo "$ICON_WARN Reverse — anime leak: $REVERSE_ANIME_COUNT (in ${SONARR_ANIME_ROOT:-<unconfigured>}, no anime signal — review, may be deliberate style placement)"
|
||||
echo "$ICON_WARN Reverse — kids leak: $REVERSE_KIDS_COUNT (in ${SONARR_KIDS_ROOT:-<unconfigured>}, adult-rated content)"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
[[ "$ENABLE_LOGGING" != true ]] && echo " (run with --log for the per-title list)"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Move Mode ━━━
|
||||
# ==============================================================================================
|
||||
# See MOVE MODE in the header for scope (forward + reverse-kids-leak, not reverse-anime-leak).
|
||||
if [[ "$MOVE_MODE" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Move Mode ━━━"
|
||||
|
||||
acquire_lock "wait"
|
||||
trap "_release_all_locks" EXIT
|
||||
|
||||
build_arr_path_map "SONARR"
|
||||
|
||||
MOVE_TARGETS=$(echo "$RESULTS" | jq -c '[.[] | select(.forward_anime_miss or .forward_kids_miss or .reverse_kids_leak)]')
|
||||
MOVE_COUNT=$(echo "$MOVE_TARGETS" | jq 'length')
|
||||
|
||||
if [[ "$MOVE_COUNT" -eq 0 ]]; then
|
||||
info "Nothing to move"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
warn "About to process $MOVE_COUNT series — one at a time, verifying after each"
|
||||
|
||||
MOVED=0
|
||||
RELOCATED_SEARCH=0
|
||||
FAILED=0
|
||||
|
||||
while IFS= read -r item; do
|
||||
id=$(echo "$item" | jq -r '.id')
|
||||
title=$(echo "$item" | jq -r '.title')
|
||||
is_anime_flag=$(echo "$item" | jq -r '.is_anime')
|
||||
is_forward_kids=$(echo "$item" | jq -r '.forward_kids_miss')
|
||||
had_files_count=$(echo "$item" | jq -r '.episodeFileCount')
|
||||
|
||||
if [[ "$is_anime_flag" == "true" ]]; then
|
||||
target_root="$SONARR_ANIME_ROOT"
|
||||
elif [[ "$is_forward_kids" == "true" ]]; then
|
||||
target_root="$SONARR_KIDS_ROOT"
|
||||
else
|
||||
target_root="$SONARR_GENERAL_ROOT"
|
||||
fi
|
||||
|
||||
if [[ -z "$target_root" ]]; then
|
||||
error " ✗ $title — target root not configured (SONARR_GENERAL_ROOT blank), skipping"
|
||||
(( FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# RESULTS only carries the reduced report fields — Sonarr's PUT expects the complete
|
||||
# resource representation, so fetch a fresh full series record to modify and send back.
|
||||
full_series=$(arr_api "$SONARR_URL" "$SONARR_API_KEY" "v3" "series/$id" "Sonarr")
|
||||
if [[ -z "$full_series" ]]; then
|
||||
error " ✗ $title — could not fetch full series record, skipping"
|
||||
(( FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
old_path=$(echo "$full_series" | jq -r '.path')
|
||||
folder_name="${old_path##*/}"
|
||||
|
||||
# A literal "/" in the folder name would build a broken nested directory instead of
|
||||
# moving to one clean folder — this is exactly the self-inflicted bug hit doing the
|
||||
# Fate/Zero and Fate/Stay Night moves by hand earlier this session.
|
||||
if [[ "$folder_name" == *"/"* ]]; then
|
||||
error " ✗ $title — folder name contains '/', skipping (needs manual handling)"
|
||||
(( FAILED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
new_path="${target_root}/${folder_name}"
|
||||
|
||||
if [[ "$had_files_count" -gt 0 ]]; then
|
||||
info " → $title: $old_path → $new_path (moving $had_files_count episode file(s))"
|
||||
move_qs="?moveFiles=true"
|
||||
else
|
||||
info " → $title: $old_path → $new_path (no files — relocating + search)"
|
||||
move_qs=""
|
||||
fi
|
||||
|
||||
updated_series=$(echo "$full_series" | jq --arg root "$target_root" --arg path "$new_path" \
|
||||
'.rootFolderPath = $root | .path = $path')
|
||||
|
||||
http_code=$(curl -sf -o /dev/null -w "%{http_code}" -X PUT \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $SONARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$updated_series" \
|
||||
"${SONARR_URL}/api/v3/series/${id}${move_qs}" 2>/dev/null)
|
||||
|
||||
if [[ "$http_code" != "200" && "$http_code" != "202" ]]; then
|
||||
error " ✗ $title — API returned HTTP $http_code — stopping (review before re-running)"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
# moveFiles=true flips rootFolderPath/episodeFileCount in the DB instantly, but the actual
|
||||
# physical move is a separate async MoveSeries command that Sonarr drains one at a time
|
||||
# internally — confirmed live: "Full House" showed episodeFileCount:192 at the new path via
|
||||
# API while the real 75GB/192 files were still sitting at the old path, MoveSeries queued
|
||||
# behind ~20 others. The DB-field check below cannot see that: poll the actual command to
|
||||
# completion first, or a batch run can report every series "moved" while most are still
|
||||
# mid-drain.
|
||||
if [[ -n "$move_qs" ]]; then
|
||||
move_cmd_id=""
|
||||
for _ in 1 2 3 4 5; do
|
||||
move_cmd_id=$(curl -sf --max-time 10 -H "X-Api-Key: $SONARR_API_KEY" \
|
||||
"${SONARR_URL}/api/v3/command" 2>/dev/null | \
|
||||
jq -r --argjson sid "$id" \
|
||||
'[.[] | select(.name == "MoveSeries" and .body.seriesId == $sid)] | sort_by(.id) | last | .id // empty' \
|
||||
2>/dev/null)
|
||||
[[ -n "$move_cmd_id" ]] && break
|
||||
sleep 1
|
||||
done
|
||||
|
||||
if [[ -z "$move_cmd_id" ]]; then
|
||||
error " ✗ $title — could not locate the MoveSeries command — stopping (review before re-running)"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
info " → $title: MoveSeries command $move_cmd_id queued, waiting for completion..."
|
||||
move_status="" move_polled=0
|
||||
while [[ "$move_polled" -lt "$SONARR_MOVE_POLL_TIMEOUT" ]]; do
|
||||
move_status=$(curl -sf --max-time 10 -H "X-Api-Key: $SONARR_API_KEY" \
|
||||
"${SONARR_URL}/api/v3/command/${move_cmd_id}" 2>/dev/null | \
|
||||
jq -r '.status // empty' 2>/dev/null)
|
||||
[[ "$move_status" == "completed" || "$move_status" == "failed" ]] && break
|
||||
sleep 10
|
||||
(( move_polled += 10 ))
|
||||
[[ $(( move_polled % 60 )) -eq 0 ]] && log " still moving $title... (${move_polled}s elapsed)"
|
||||
done
|
||||
|
||||
if [[ "$move_status" != "completed" ]]; then
|
||||
error " ✗ $title — MoveSeries command $move_cmd_id ended as '${move_status:-timed out after ${SONARR_MOVE_POLL_TIMEOUT}s}' — stopping"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
fi
|
||||
|
||||
sleep 3
|
||||
|
||||
# Never trust the PUT response alone — re-fetch and confirm the change actually landed.
|
||||
# This exact check is what caught the earlier race condition doing this by hand: two
|
||||
# series reported "success" while episodeFileCount had silently dropped to 0.
|
||||
verify_series=$(arr_api "$SONARR_URL" "$SONARR_API_KEY" "v3" "series/$id" "Sonarr")
|
||||
verify_root=$(echo "$verify_series" | jq -r '.rootFolderPath')
|
||||
verify_filecount=$(echo "$verify_series" | jq -r '.statistics.episodeFileCount // 0')
|
||||
|
||||
if [[ "$verify_root" != "$target_root" ]]; then
|
||||
error " ✗ $title — verification failed (root: $verify_root) — stopping"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
|
||||
if [[ "$had_files_count" -gt 0 ]]; then
|
||||
if [[ "$verify_filecount" -eq "$had_files_count" ]]; then
|
||||
echo " $ICON_SUCCESS $title — moved and verified ($verify_filecount files)"
|
||||
(( MOVED++ ))
|
||||
else
|
||||
error " ✗ $title — verification failed (root updated but episode count $verify_filecount != expected $had_files_count) — stopping"
|
||||
(( FAILED++ ))
|
||||
break
|
||||
fi
|
||||
else
|
||||
search_code=$(curl -sf -o /dev/null -w "%{http_code}" -X POST \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $SONARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "{\"name\":\"SeriesSearch\",\"seriesId\":${id}}" \
|
||||
"${SONARR_URL}/api/v3/command" 2>/dev/null)
|
||||
if [[ "$search_code" == "200" || "$search_code" == "201" ]]; then
|
||||
echo " $ICON_SUCCESS $title — relocated, search triggered"
|
||||
else
|
||||
warn " $title — relocated but search trigger returned HTTP $search_code (will pick up on next scheduled search)"
|
||||
fi
|
||||
(( RELOCATED_SEARCH++ ))
|
||||
fi
|
||||
done < <(echo "$MOVE_TARGETS" | jq -c '.[]')
|
||||
|
||||
# arr_get_tracked_data() is cache-first — every write above changed rootFolderPath, so the
|
||||
# shared cache is now stale until the next scheduled arr_cache_prefill run. Refresh it now
|
||||
# rather than leave that window open for every other script reading this cache.
|
||||
if [[ "$(( MOVED + RELOCATED_SEARCH ))" -gt 0 ]]; then
|
||||
info "Refreshing shared tracked-data cache..."
|
||||
fresh_series=$(arr_api "$SONARR_URL" "$SONARR_API_KEY" "v3" "series" "Sonarr")
|
||||
[[ -n "$fresh_series" ]] && arr_cache_write "sonarr" "$fresh_series"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY MOVE SUMMARY ━━━━━"
|
||||
echo "$ICON_SUCCESS Moved (files relocated): $MOVED"
|
||||
echo "$ICON_SUCCESS Relocated + search triggered: $RELOCATED_SEARCH"
|
||||
echo "$ICON_ERROR Failed: $FAILED"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
fi
|
||||
|
||||
exit 0
|
||||
@@ -0,0 +1,619 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Sonarr Cleanup =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Delete orphaned TV episode files not tracked by Sonarr. Queries the API for
|
||||
# all tracked episode file paths, walks the library on disk, and removes anything
|
||||
# untracked that is old enough to be past the import window. Triggers an Emby
|
||||
# library clean after each deletion run so ghost entries disappear immediately.
|
||||
#
|
||||
# The tracked-count floor check (Safety Layer 6) is rescan-aware: if Sonarr's own
|
||||
# RescanSeries/DownloadedEpisodesScan is active (independently of this script's own
|
||||
# lighter ProcessMonitoredDownloads pre-flight — e.g. during a large missing-episode
|
||||
# search campaign), a genuinely low mid-scan count gets waited out (calibrated to that
|
||||
# command's historical duration via arr_get_rescan_duration(), up to 3 strikes) and
|
||||
# re-fetched rather than triggering a false-alarm abort. Mirrors the same fix built for
|
||||
# lidarr_cleanup.sh 2026-07-16 after a whole-library rescan there made trackFileCount
|
||||
# read 22% of normal mid-scan.
|
||||
#
|
||||
# Cache-first, both layers (2026-07-17). The series list itself comes from the shared
|
||||
# tracked-data cache via arr_get_tracked_data() — fresh (kept warm every 30min by
|
||||
# arr_cache_prefill.sh), live fetch as fallback. The per-series episodefile walk below still
|
||||
# always fetches live (that's the actual disk-truth this script's delete decisions depend
|
||||
# on), but write-throughs its result to arr_item_cache_write() for any future script that
|
||||
# needs Sonarr's per-episode data — no second consumer exists yet, unlike Lidarr's
|
||||
# lidarr_missing_art.sh, but the data's there once one does. The filesystem is walked once
|
||||
# per run, not twice — classification records which paths are eligible for deletion as it
|
||||
# goes, and the delete pass (once the size-threshold check below passes) just acts on that
|
||||
# list instead of re-walking and re-classifying the whole tree. That single walk also gets
|
||||
# size+ctime straight from find -printf instead of a separate stat fork per file — find
|
||||
# already has to stat() every entry to know it's -type f, so this is free by comparison.
|
||||
# Measured ~130x faster per file (0.033ms vs 4.3ms).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Every file encountered on disk is classified into one of five categories:
|
||||
#
|
||||
# TRACKED — Sonarr API knows this exact path → leave it alone
|
||||
# PROTECTED — matches SONARR_PROTECTED_PATTERNS → never delete
|
||||
# ORPHAN — video file, not tracked, older than SONARR_ORPHAN_AGE → delete
|
||||
# JUNK — not a video extension, not protected → delete regardless of age
|
||||
# RECENT — not tracked, under SONARR_ORPHAN_AGE → skip (may be mid-import)
|
||||
#
|
||||
# Sonarr generates show artwork (*.jpg), metadata (*.nfo), and manages subtitles
|
||||
# (*.srt, *.sub, *.ass) but does NOT include these in its tracked file API response.
|
||||
# Without PROTECTED classification these would be deleted — breaking Sonarr and
|
||||
# Emby metadata display.
|
||||
#
|
||||
# After deletions: notify_emby_scan() triggers Emby "Clean Missing Files" task.
|
||||
# Emby removes ghost entries immediately — no user-facing file-not-found errors.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# API as Ground Truth
|
||||
# What Sonarr tracks is authoritative. Files not in the API response are
|
||||
# orphans — Sonarr has no record of them and they serve no purpose.
|
||||
# The script never infers ownership from directory structure alone.
|
||||
#
|
||||
# Age Gate Before Deletion
|
||||
# Files under SONARR_ORPHAN_AGE are left alone regardless of tracked status.
|
||||
# Sonarr's import pipeline writes files before registering them — acting
|
||||
# immediately would delete files mid-import.
|
||||
# Age is measured from ctime, not mtime — an import preserves the release's original
|
||||
# mtime, so a file that landed today can read as years old and skip this gate. Depends
|
||||
# on media_shares_permissions.sh touching only entries that are actually wrong.
|
||||
#
|
||||
# Emby Cleanup Is Part of the Job
|
||||
# Deleting a file without telling Emby leaves ghost entries that show as
|
||||
# broken items. Triggering the Emby clean is not optional — it completes
|
||||
# the deletion from the user's perspective.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Seven gates — ALL must pass before any file is touched:
|
||||
# 1. Container running and not starting/unhealthy
|
||||
# 2. API reachable
|
||||
# 3. API version matches SONARR_VERSION_MAJOR in master.conf
|
||||
# 4. Series count > 0
|
||||
# 5. Tracked file count > 0
|
||||
# 6. Tracked count >= SONARR_MIN_TRACKED_PCT % of last known count
|
||||
# 7. Deletion size < SONARR_MAX_DELETE_GB — or --i-know-what-im-doing required
|
||||
#
|
||||
# acquire_lock "wait" — large scans take time, wait for previous run to finish
|
||||
# jq + curl validation — exits if either tool missing
|
||||
# ARR_DOCKER_TIMEOUT — container checks protected against daemon hangs (script-local, not common.sh's DOCKER_TIMEOUT)
|
||||
# notify_emby_scan() — triggers Emby clean after deletion
|
||||
# Silent by default — orphans/junk warn(), clean library logs silently
|
||||
#
|
||||
# ==============================================================================================
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# SONARR_TRACKED_COUNT_FILE — persistent baseline for the tracked % safety check (gate 6)
|
||||
# Updated after each successful run. Protects against misconfigured root path
|
||||
# returning an empty API response and deleting the entire library.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_SONARR_URL / HOST*_SONARR_API_KEY / HOST*_SONARR_TV_ROOT
|
||||
# HOST*_SONARR_PATH_MAP — container path → host path translation
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# SONARR_ORPHAN_AGE — days before untracked file eligible for deletion
|
||||
# SONARR_MAX_DELETE_GB — require --i-know-what-im-doing above this
|
||||
# SONARR_MIN_TRACKED_PCT — abort if tracked count drops below this % of last run
|
||||
# SONARR_TRACKED_COUNT_FILE — persistent baseline file path
|
||||
# SONARR_EXTENSIONS — video file extensions for orphan classification
|
||||
# SONARR_PROTECTED_PATTERNS — file patterns never deleted
|
||||
# SONARR_VERSION_MAJOR — expected Sonarr major version for API safety check
|
||||
# SONARR_IMPORT_SCAN_TIMEOUT — seconds to wait for pre-flight import scan (default 600)
|
||||
# ARR_CLEANUP_STATS — stats file path (read by coffee report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# sonarr_cleanup.sh — normal run
|
||||
# sonarr_cleanup.sh --dry-run — preview, no deletions
|
||||
# sonarr_cleanup.sh --log — verbose output
|
||||
# sonarr_cleanup.sh --status — show config and exit
|
||||
# sonarr_cleanup.sh --i-know-what-im-doing — bypass size threshold
|
||||
# sonarr_cleanup.sh --i-know-what-im-doing --skip-age-check — NUCLEAR MODE
|
||||
#
|
||||
# NUCLEAR MODE: both flags bypass age check AND size threshold. User accepts full
|
||||
# responsibility — the flag name is long and annoying by design.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# ── Special flag pre-processing ───────────────────────────────────────────────────────────────
|
||||
parse_destructive_flags "$@"
|
||||
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
|
||||
# ── Nuclear mode warning ──────────────────────────────────────────────────────────────────────
|
||||
nuclear_mode_warning
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for Sonarr API calls"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
notify "Sonarr cleanup failed on $(hostname) — jq not installed" "Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
||||
acquire_lock "wait"
|
||||
TMP_DIR="/tmp/sonarr_cleanup_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "_release_all_locks; rm -rf $TMP_DIR" EXIT
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases SONARR_URL, SONARR_API_KEY, SONARR_TV_ROOT
|
||||
detect_hosts
|
||||
|
||||
# Skip if Sonarr is not configured on this host
|
||||
if [[ -z "${SONARR_URL:-}" ]] || [[ -z "${SONARR_API_KEY:-}" ]]; then
|
||||
info "Sonarr not configured on $MY_ID ($LOCAL_SERVER_NAME) — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ARR_DOCKER_TIMEOUT=15
|
||||
SONARR_CONTAINER="Sonarr"
|
||||
|
||||
# Build path map from MY_ID's Sonarr path map
|
||||
build_arr_path_map "SONARR"
|
||||
|
||||
require_var SONARR_URL
|
||||
require_var SONARR_API_KEY
|
||||
require_var SONARR_TV_ROOT
|
||||
|
||||
if [[ ! -d "$SONARR_TV_ROOT" ]]; then
|
||||
error "TV root not found: $SONARR_TV_ROOT"
|
||||
notify "Sonarr cleanup failed on $(hostname) — TV root not found: $SONARR_TV_ROOT" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Config: url=${SONARR_URL} root=${SONARR_TV_ROOT}"
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $SONARR_URL"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be deleted"
|
||||
[[ "$I_KNOW" == true ]] && warn "OVERRIDE — --i-know-what-im-doing active"
|
||||
[[ "$SKIP_AGE_CHECK" == true ]] && warn "OVERRIDE — --skip-age-check active — age check bypassed"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Sonarr URL: $SONARR_URL"
|
||||
echo "$ICON_GEAR TV root: $SONARR_TV_ROOT"
|
||||
echo "$ICON_TIME Orphan age: ${SONARR_ORPHAN_AGE} days"
|
||||
echo "$ICON_GEAR Max delete: ${SONARR_MAX_DELETE_GB}GB (requires --i-know-what-im-doing)"
|
||||
echo "$ICON_GEAR Min tracked %: ${SONARR_MIN_TRACKED_PCT}%"
|
||||
echo "$ICON_GEAR Sonarr ver: v${SONARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Extensions: ${SONARR_EXTENSIONS[*]}"
|
||||
echo "$ICON_GEAR Protected patterns: ${SONARR_PROTECTED_PATTERNS[*]}"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "$ICON_GEAR I know: $I_KNOW"
|
||||
echo "$ICON_GEAR Skip age check: $SKIP_AGE_CHECK"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 1 — Container Health ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
|
||||
check_container_health "$SONARR_CONTAINER" "$ARR_DOCKER_TIMEOUT" "Sonarr Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# check_container_health(), arr_api(), has_extension(), matches_pattern_list(), format_bytes() — common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Pre-flight: Sonarr Import Scan ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Pre-flight: Sonarr Import Scan ━━━"
|
||||
|
||||
# Fetch root folders from Sonarr API and translate container paths to host paths
|
||||
mapfile -t SCAN_ROOTS < <(
|
||||
arr_api "$SONARR_URL" "$SONARR_API_KEY" "v3" "rootfolder" "Sonarr" | \
|
||||
jq -r '.[].path' 2>/dev/null | \
|
||||
while IFS= read -r cp; do translate_path "$cp"; done
|
||||
)
|
||||
|
||||
if [[ "${#SCAN_ROOTS[@]}" -eq 0 ]]; then
|
||||
error "No root folders returned from Sonarr API — aborting"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — no root folders from API" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "Scan targets (${#SCAN_ROOTS[@]}): ${SCAN_ROOTS[*]}"
|
||||
|
||||
info "Triggering ProcessMonitoredDownloads pre-flight"
|
||||
SCAN_PAYLOAD='{"name": "ProcessMonitoredDownloads"}'
|
||||
|
||||
trigger_and_await_command "$SONARR_URL" "$SONARR_API_KEY" "v3" "$SCAN_PAYLOAD" "${SONARR_IMPORT_SCAN_TIMEOUT:-600}" "sonarr"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Sonarr Tracked Files ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching Sonarr Tracked Files ━━━"
|
||||
|
||||
# Safety Layer 2 — API reachability
|
||||
if ! check_api "$SONARR_URL" "Sonarr" 10; then
|
||||
notify "Sonarr cleanup aborted on $(hostname) — API unreachable" "Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Safety Layer 3 — API version check
|
||||
check_arr_version "$SONARR_URL" "$SONARR_API_KEY" "v3" "$SONARR_VERSION_MAJOR" "Sonarr" || exit 1
|
||||
|
||||
info "Querying Sonarr API..."
|
||||
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (now kept
|
||||
# current every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to
|
||||
# a live fetch when it's stale, and waits out an active rescan before either. Only the series
|
||||
# list itself is cached — the per-series episodeFile data below is never cached and always
|
||||
# live, since that's the actual disk-truth this script's cleanup decisions depend on.
|
||||
SERIES_RESPONSE=$(arr_get_tracked_data "sonarr" "$SONARR_URL" "$SONARR_API_KEY" "v3") || {
|
||||
error "Failed to fetch series from Sonarr"
|
||||
notify "Sonarr cleanup failed on $(hostname) — could not fetch series" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
}
|
||||
|
||||
SERIES_IDS=$(echo "$SERIES_RESPONSE" | jq -r '.[].id' 2>/dev/null)
|
||||
SERIES_COUNT=$(echo "$SERIES_IDS" | grep -c "." 2>/dev/null || true)
|
||||
|
||||
# Safety Layer 4 — series count > 0
|
||||
if [[ "$SERIES_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 series — aborting to prevent mass deletion"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — 0 series returned" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "Found $SERIES_COUNT series — fetching episode files..."
|
||||
|
||||
TRACKED_FILE="$TMP_DIR/tracked_paths.txt"
|
||||
> "$TRACKED_FILE"
|
||||
|
||||
# Fetches every series' episode-file paths fresh into TRACKED_FILE/TRACKED_MAP/TRACKED_COUNT.
|
||||
# Pulled into a function so the rescan-aware retry below can re-fetch after waiting without
|
||||
# duplicating this whole loop inline.
|
||||
#
|
||||
# Write-through — also caches the raw per-episode data via arr_item_cache_write() (2026-07-17)
|
||||
# for any future script that needs Sonarr's per-episode file data — none exist yet (unlike
|
||||
# Lidarr, where lidarr_missing_art.sh already reads this), but the walk is happening regardless
|
||||
# for our own cleanup decisions, so the cache write is free by comparison. Future consumers:
|
||||
# call arr_get_cached_items("sonarr") first, fall back to your own live walk on a miss.
|
||||
_fetch_tracked_files() {
|
||||
> "$TRACKED_FILE"
|
||||
local _series_index=0
|
||||
local all_episodes_tmp
|
||||
all_episodes_tmp=$(mktemp)
|
||||
while IFS= read -r series_id; do
|
||||
[[ -z "$series_id" ]] && continue
|
||||
(( _series_index++ ))
|
||||
[[ $(( _series_index % 50 )) -eq 0 ]] && \
|
||||
log "Fetching files: $_series_index/$SERIES_COUNT series..."
|
||||
SERIES_FILES=$(arr_api "$SONARR_URL" "$SONARR_API_KEY" "v3" "episodefile?seriesId=${series_id}" "Sonarr" 2>/dev/null)
|
||||
if [[ -n "$SERIES_FILES" ]]; then
|
||||
echo "$SERIES_FILES" >> "$all_episodes_tmp"
|
||||
while IFS= read -r api_path; do
|
||||
[[ -z "$api_path" ]] && continue
|
||||
translate_path "$api_path" >> "$TRACKED_FILE"
|
||||
done < <(echo "$SERIES_FILES" | jq -r '.[].path' 2>/dev/null)
|
||||
fi
|
||||
done <<< "$SERIES_IDS"
|
||||
|
||||
arr_item_cache_write "sonarr" "$(jq -s 'add // []' "$all_episodes_tmp" 2>/dev/null)"
|
||||
rm -f "$all_episodes_tmp"
|
||||
|
||||
sort -u "$TRACKED_FILE" -o "$TRACKED_FILE"
|
||||
|
||||
# Build in-memory lookup map — O(1) per lookup vs O(n) grep per file
|
||||
unset TRACKED_MAP
|
||||
declare -gA TRACKED_MAP
|
||||
while IFS= read -r _tracked_path; do
|
||||
[[ -n "$_tracked_path" ]] && TRACKED_MAP["$_tracked_path"]=1
|
||||
done < "$TRACKED_FILE"
|
||||
unset _tracked_path
|
||||
TRACKED_COUNT=$(wc -l < "$TRACKED_FILE")
|
||||
}
|
||||
|
||||
_fetch_tracked_files
|
||||
info "Built in-memory lookup map: ${#TRACKED_MAP[@]} tracked paths"
|
||||
|
||||
# Safety Layer 5 — tracked count > 0
|
||||
if [[ "$TRACKED_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 tracked files — aborting to prevent mass deletion"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — 0 tracked files returned" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "$SERIES_COUNT series | $TRACKED_COUNT tracked episode files"
|
||||
|
||||
# Safety Layer 6 — percentage drop vs last known count, with rescan-aware retry.
|
||||
# ProcessMonitoredDownloads (this script's own pre-flight) is a different, lighter operation
|
||||
# than a full library rescan — but Sonarr's own RescanSeries/DownloadedEpisodesScan can be
|
||||
# triggered independently (e.g. during a large missing-episode search campaign) and would
|
||||
# cause the exact same mid-scan count dip confirmed on Lidarr 2026-07-16. Wait it out
|
||||
# (calibrated to that command's own historical duration) before treating a drop as genuine.
|
||||
_last_known=$(cat "$SONARR_TRACKED_COUNT_FILE" 2>/dev/null || echo 0)
|
||||
if [[ "$_last_known" -gt 0 ]]; then
|
||||
_strike=1
|
||||
while [[ "$_strike" -le 3 ]]; do
|
||||
_pct=$(awk "BEGIN {printf \"%d\", ($TRACKED_COUNT / $_last_known) * 100}")
|
||||
[[ "$_pct" -ge "${SONARR_MIN_TRACKED_PCT:-50}" ]] && break
|
||||
|
||||
_active_cmd=$(arr_active_rescan_command "sonarr" "$SONARR_URL" "$SONARR_API_KEY" "v3")
|
||||
[[ -z "$_active_cmd" ]] && break # low count, nothing rescanning — genuine, don't retry
|
||||
|
||||
_wait=$(( $(arr_get_rescan_duration "sonarr" "$_active_cmd" 300) / 2 ))
|
||||
[[ "$_wait" -lt 30 ]] && _wait=30
|
||||
warn "Tracked count ${_pct}% of last run, but $_active_cmd active — waiting ${_wait}s (strike ${_strike}/3)"
|
||||
sleep "$_wait"
|
||||
_fetch_tracked_files
|
||||
(( _strike++ ))
|
||||
done
|
||||
|
||||
if [[ "$_strike" -gt 3 ]]; then
|
||||
_active_cmd=$(arr_active_rescan_command "sonarr" "$SONARR_URL" "$SONARR_API_KEY" "v3")
|
||||
if [[ -n "$_active_cmd" ]]; then
|
||||
warn "Sonarr still busy ($_active_cmd) after 3 strikes — deferring to next scheduled run"
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
check_tracked_count_floor "$TRACKED_COUNT" "$SONARR_TRACKED_COUNT_FILE" "$SONARR_MIN_TRACKED_PCT" "Sonarr Cleanup"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Scan TV Root ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Scanning TV Root ━━━"
|
||||
info "Root: $SONARR_TV_ROOT | Orphan age: ${SONARR_ORPHAN_AGE} days"
|
||||
|
||||
START=$(date +%s)
|
||||
ORPHAN_COUNT=0
|
||||
JUNK_COUNT=0
|
||||
RECENT_COUNT=0
|
||||
PROTECTED_COUNT=0
|
||||
ORPHAN_BYTES=0
|
||||
JUNK_BYTES=0
|
||||
|
||||
AGE_SECONDS=$(( SONARR_ORPHAN_AGE * 86400 ))
|
||||
NOW=$(date +%s)
|
||||
|
||||
# Files classified ORPHAN/JUNK below get their path recorded here, so the deletion pass can
|
||||
# just delete them directly instead of re-walking and re-classifying every SCAN_ROOTS entry a
|
||||
# second time (2026-07-17) — the size-threshold check below needs to know the total before
|
||||
# deleting anything, not before knowing what to delete.
|
||||
TO_DELETE_FILE="$TMP_DIR/to_delete_paths.txt"
|
||||
> "$TO_DELETE_FILE"
|
||||
|
||||
# ── Orphan strikes ────────────────────────────────────────────────────────────────────────────
|
||||
# Same contract as radarr_cleanup.sh: a file must classify for deletion on
|
||||
# SONARR_ORPHAN_STRIKE_LIMIT consecutive runs before it is removed. Covers the partial
|
||||
# classification failure that is too small to trip the tracked-count floor above. The file is
|
||||
# rebuilt from each run rather than edited, which is what prunes it.
|
||||
SONARR_ORPHAN_STRIKE_LIMIT="${SONARR_ORPHAN_STRIKE_LIMIT:-2}"
|
||||
STRIKES_FILE="${SONARR_ORPHAN_STRIKES_FILE:-$DB_DIR/sonarr_orphan_strikes.tsv}"
|
||||
mkdir -p "$(dirname "$STRIKES_FILE")" 2>/dev/null || true
|
||||
touch "$STRIKES_FILE" 2>/dev/null || true
|
||||
STRIKES_NEW="$TMP_DIR/strikes_new.tsv"
|
||||
> "$STRIKES_NEW"
|
||||
HELD_COUNT=0
|
||||
HELD_BYTES=0
|
||||
|
||||
orphan_strike_ok() {
|
||||
local path="$1" prev strikes
|
||||
prev=$(wd_state_get "$path" "$STRIKES_FILE"); prev="${prev//[^0-9]/}"
|
||||
strikes=$(( ${prev:-0} + 1 ))
|
||||
printf '%s:%s\n' "$path" "$strikes" >> "$STRIKES_NEW"
|
||||
(( strikes >= SONARR_ORPHAN_STRIKE_LIMIT )) && return 0
|
||||
warn " strike $strikes/$SONARR_ORPHAN_STRIKE_LIMIT — not removing yet: $path"
|
||||
return 1
|
||||
}
|
||||
|
||||
while read -r FILE_SIZE FILE_CTIME filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
FILE_CTIME="${FILE_CTIME%%.*}"
|
||||
|
||||
if [[ -n "${TRACKED_MAP[$filepath]:-}" ]]; then
|
||||
log "TRACKED: $filepath"
|
||||
continue
|
||||
fi
|
||||
|
||||
if matches_pattern_list "$filepath" "${SONARR_PROTECTED_PATTERNS[@]}"; then
|
||||
log "$ICON_PROTECTED PROTECTED: $filepath"
|
||||
(( PROTECTED_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
if has_extension "$filepath" "${SONARR_EXTENSIONS[@]}"; then
|
||||
# ctime, not mtime — an import preserves the release's original mtime, so a file
|
||||
# Sonarr moved in today can read as years old and skip this gate entirely.
|
||||
# Measured 2026-07-27: 400 of 400 files imported that week had mtimes over 7
|
||||
# days, one of them 9613 days. ctime is stamped when the file lands on this
|
||||
# filesystem and cannot be carried in from an archive. This only holds because
|
||||
# media_shares_permissions.sh applies owner/mode conditionally — a blanket
|
||||
# chown/chmod restamps every inode nightly and would peg every file at age 0.
|
||||
FILE_AGE=$(( NOW - FILE_CTIME ))
|
||||
|
||||
if [[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && [[ "$SKIP_AGE_CHECK" != true ]]; then
|
||||
log "RECENT (skipping): $filepath"
|
||||
(( RECENT_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
warn "$ICON_TRASH ORPHAN: $filepath"
|
||||
(( ORPHAN_COUNT++ ))
|
||||
ORPHAN_BYTES=$(( ORPHAN_BYTES + FILE_SIZE ))
|
||||
if ! orphan_strike_ok "$filepath"; then (( HELD_COUNT++ )); HELD_BYTES=$(( HELD_BYTES + FILE_SIZE )); continue; fi
|
||||
printf '%s\t%s\t%s\n' "$FILE_SIZE" "$FILE_CTIME" "$filepath" >> "$TO_DELETE_FILE"
|
||||
else
|
||||
log "JUNK: $filepath"
|
||||
(( JUNK_COUNT++ ))
|
||||
JUNK_BYTES=$(( JUNK_BYTES + FILE_SIZE ))
|
||||
if ! orphan_strike_ok "$filepath"; then (( HELD_COUNT++ )); HELD_BYTES=$(( HELD_BYTES + FILE_SIZE )); continue; fi
|
||||
printf '%s\t%s\t%s\n' "$FILE_SIZE" "$FILE_CTIME" "$filepath" >> "$TO_DELETE_FILE"
|
||||
fi
|
||||
|
||||
# -printf gets size + mtime directly from find's own stat() during the walk, instead of a
|
||||
# separate stat fork per file (2026-07-17) — measured ~130x faster per file (0.033ms vs
|
||||
# 4.3ms), since find already has to stat() every entry anyway to know it's -type f.
|
||||
done < <(
|
||||
for host_path in "${SCAN_ROOTS[@]}"; do
|
||||
[[ -d "$host_path" ]] && find "$host_path" -type f -printf '%s %C@ %p\n' 2>/dev/null
|
||||
done | sort -u
|
||||
)
|
||||
|
||||
# Eligible, not classified: a file still serving its strikes is an orphan but is not queued this
|
||||
# run, so it must not appear in the denominator the budget reports against.
|
||||
TOTAL_DELETE_BYTES=$(( ORPHAN_BYTES + JUNK_BYTES - HELD_BYTES ))
|
||||
TOTAL_REMOVED=$(( ORPHAN_COUNT + JUNK_COUNT - HELD_COUNT ))
|
||||
|
||||
# Rebuilt, never edited. Skipped on a dry run: a preview that advanced real counters would make
|
||||
# the next real run delete a run early.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
mv "$STRIKES_NEW" "$STRIKES_FILE" 2>/dev/null || warn "Could not update $STRIKES_FILE"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 7 — Deletion Size Threshold ━━━
|
||||
# ==============================================================================================
|
||||
# A per-run budget, not a veto — see apply_delete_budget() in common.sh. The ceiling still caps
|
||||
# any single run; it just no longer deadlocks on a backlog larger than itself.
|
||||
BUDGET_FILE="$TMP_DIR/to_delete_budgeted.txt"
|
||||
|
||||
if [[ "$I_KNOW" == true ]]; then
|
||||
warn "OVERRIDE — --i-know-what-im-doing active, per-run budget not applied"
|
||||
cut -d"$(printf '\t')" -f3- "$TO_DELETE_FILE" > "$BUDGET_FILE"
|
||||
_BUDGET_KEPT_COUNT=$TOTAL_REMOVED; _BUDGET_KEPT_BYTES=$TOTAL_DELETE_BYTES
|
||||
_BUDGET_DEFERRED_COUNT=0; _BUDGET_DEFERRED_BYTES=0; _BUDGET_STUCK=""
|
||||
else
|
||||
apply_delete_budget "$TO_DELETE_FILE" "$BUDGET_FILE" "$SONARR_MAX_DELETE_GB"
|
||||
if [[ -n "$_BUDGET_STUCK" ]]; then
|
||||
error "Single file exceeds the ${SONARR_MAX_DELETE_GB}GB budget on its own — nothing removed this run"
|
||||
error " $_BUDGET_STUCK"
|
||||
error "Raise SONARR_MAX_DELETE_GB or clear this one with --i-know-what-im-doing"
|
||||
notify "Sonarr cleanup stalled on $(hostname) — one file exceeds the ${SONARR_MAX_DELETE_GB}GB budget" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
elif [[ "$_BUDGET_DEFERRED_COUNT" -gt 0 ]]; then
|
||||
warn "Budget ${SONARR_MAX_DELETE_GB}GB — removing $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED ($(format_bytes "$_BUDGET_KEPT_BYTES")), deferring $_BUDGET_DEFERRED_COUNT ($(format_bytes "$_BUDGET_DEFERRED_BYTES")) to the next run"
|
||||
notify "Sonarr cleanup removed $(format_bytes "$_BUDGET_KEPT_BYTES") of $(format_bytes "$TOTAL_DELETE_BYTES") on $(hostname) — $_BUDGET_DEFERRED_COUNT file(s) deferred" \
|
||||
"Sonarr Cleanup" "normal"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Execute Deletions ─────────────────────────────────────────────────────────────────────────
|
||||
# Reuses TO_DELETE_FILE from the classification pass above instead of re-walking and
|
||||
# re-classifying every SCAN_ROOTS entry again.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
rm -f "$filepath" 2>/dev/null || error "Failed to delete: $filepath"
|
||||
done < "$BUDGET_FILE"
|
||||
|
||||
info "Cleaning up empty folders..."
|
||||
for host_path in "${SCAN_ROOTS[@]}"; do
|
||||
[[ -d "$host_path" ]] && \
|
||||
find "$host_path" -mindepth 1 -type d -empty -delete 2>/dev/null
|
||||
done
|
||||
info "Empty folders removed"
|
||||
fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
ORPHAN_HUMAN=$(format_bytes "$ORPHAN_BYTES")
|
||||
JUNK_HUMAN=$(format_bytes "$JUNK_BYTES")
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY SONARR CLEANUP SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Tracked: $TRACKED_COUNT files ($SERIES_COUNT series)"
|
||||
echo "$ICON_SHIELD Protected: $PROTECTED_COUNT files (artwork, subtitles, metadata)"
|
||||
echo "$ICON_TRASH Orphans: $ORPHAN_COUNT files ($ORPHAN_HUMAN)"
|
||||
echo "$ICON_TRASH Junk: $JUNK_COUNT files ($JUNK_HUMAN)"
|
||||
echo "$ICON_SKIP Recent skipped: $RECENT_COUNT files (under ${SONARR_ORPHAN_AGE} days)"
|
||||
[[ "${HELD_COUNT:-0}" -gt 0 ]] && \
|
||||
echo "$ICON_SKIP Held (strikes): $HELD_COUNT files ($(format_bytes "$HELD_BYTES")) — under ${SONARR_ORPHAN_STRIKE_LIMIT} consecutive runs"
|
||||
[[ "${_BUDGET_DEFERRED_COUNT:-0}" -gt 0 ]] && \
|
||||
echo "$ICON_SKIP Deferred: $_BUDGET_DEFERRED_COUNT files ($(format_bytes "$_BUDGET_DEFERRED_BYTES")) — over the ${SONARR_MAX_DELETE_GB}GB run budget"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no files deleted"
|
||||
elif [[ "$TOTAL_REMOVED" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Clean — nothing to remove"
|
||||
else
|
||||
# What was actually removed, not what was classified. With strikes and a budget in force those
|
||||
# differ, and reporting the classification as the outcome is the oldest bug shape here.
|
||||
warn "$ICON_DONE Removed $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED eligible files ($(format_bytes "$_BUDGET_KEPT_BYTES"))"
|
||||
notify "Sonarr cleanup on $(hostname) — removed $_BUDGET_KEPT_COUNT of $TOTAL_REMOVED eligible files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
# Notify Emby to clean missing files — removes ghost entries immediately
|
||||
notify_emby_scan
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# Write stats for sunday_morning_coffee_report.sh
|
||||
if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_CLEANUP_STATS:-}" ]]; then
|
||||
echo "$(date '+%Y-%m-%d')|sonarr|${ORPHAN_COUNT}|${ORPHAN_BYTES}|${JUNK_COUNT}|${JUNK_BYTES}|${RECENT_COUNT}|${TRACKED_COUNT}" \
|
||||
>> "$ARR_CLEANUP_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
Regular → Executable
+31
-5
@@ -9,6 +9,10 @@
|
||||
# status="deleted" — they generate health errors and can never be monitored
|
||||
# or downloaded.
|
||||
#
|
||||
# Cache-first series list (2026-07-17) — comes from the shared tracked-data cache via
|
||||
# arr_get_tracked_data(), fresh (kept warm every 30min by arr_cache_prefill.sh), live fetch
|
||||
# as fallback.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
@@ -23,6 +27,25 @@
|
||||
# Per-deletion output is always visible — deletions are never silently swallowed.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Health Error Hygiene
|
||||
# status="deleted" series can never be monitored or downloaded — they only
|
||||
# generate persistent health errors. Removing them is maintenance, not
|
||||
# data loss: most have no associated files.
|
||||
#
|
||||
# Conservative File Handling
|
||||
# Files are not deleted by default. A TVDB-removed series may still have
|
||||
# episodes on disk that the user wants to keep. The --delete-files flag
|
||||
# is an explicit opt-in, not the default path.
|
||||
#
|
||||
# Exclusion List Prevents Re-add
|
||||
# Removed entries are added to Sonarr's import exclusion list by default.
|
||||
# Without this, the same deleted series can be re-added by lists or searches
|
||||
# and immediately generate the same health error again.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -94,12 +117,13 @@ acquire_lock
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "$SONARR_URL" ]]; then
|
||||
log "Sonarr not configured for $MY_ID — nothing to do"
|
||||
echo "Sonarr not configured for $MY_ID — skipping"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ADD_EXCLUSION="${SONARR_DROPPED_ADD_EXCLUSION:-true}"
|
||||
|
||||
log "$ICON_GEAR Config: url=${SONARR_URL} add-exclusion=${ADD_EXCLUSION} delete-files=${DELETE_FILES:-false}"
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $SONARR_URL"
|
||||
[[ "$DELETE_FILES" == true ]] && warn "DELETE FILES MODE — files will be removed from disk"
|
||||
[[ "$DELETE_FILES" == false ]] && echo " Files: records only (use --delete-files to also remove from disk)"
|
||||
@@ -132,8 +156,10 @@ if ! curl -sf --connect-timeout 5 --max-time 10 \
|
||||
exit 1
|
||||
fi
|
||||
|
||||
SERIES=$(curl -sf --connect-timeout 5 --max-time 30 \
|
||||
"$SONARR_URL/api/v3/series?apikey=$SONARR_API_KEY" 2>/dev/null)
|
||||
# Cache-first — arr_get_tracked_data() serves the shared cache when it's fresh (kept current
|
||||
# every 30min by arr_cache_prefill.sh in CRITICAL_MAINTENANCE_SCRIPTS), falls back to a live
|
||||
# fetch when it's stale, and waits out an active rescan before either.
|
||||
SERIES=$(arr_get_tracked_data "sonarr" "$SONARR_URL" "$SONARR_API_KEY" "v3")
|
||||
|
||||
if [[ -z "$SERIES" || "$SERIES" == "null" ]]; then
|
||||
error "Sonarr series API returned empty"
|
||||
@@ -145,7 +171,7 @@ DROPPED=$(echo "$SERIES" | jq '[.[] | select(.status == "deleted")] | length')
|
||||
echo " $TOTAL series total — $DROPPED dropped from TVDB"
|
||||
|
||||
if [[ "$DROPPED" -eq 0 ]]; then
|
||||
log "No TVDB-removed series found — nothing to do"
|
||||
echo "No TVDB-removed series found — nothing to do"
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY SONARR TVDB REMOVED SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
@@ -199,7 +225,7 @@ while IFS=$'\t' read -r id title year tvdb_id episode_file_count size_on_disk; d
|
||||
CURL_EXIT=$?
|
||||
|
||||
if [[ "$CURL_EXIT" -eq 0 ]]; then
|
||||
log " Removed from Sonarr ✅"
|
||||
echo " Removed from Sonarr ✅"
|
||||
REMOVED+=("$title")
|
||||
if [[ "$DELETE_PARAM" == "true" ]]; then
|
||||
(( FILES_DELETED++ ))
|
||||
Executable
+188
@@ -0,0 +1,188 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ======================= Upgrade Webhook Listener (continuous) ================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Starts a standalone Node.js HTTP server that receives Sonarr/Radarr/Lidarr
|
||||
# OnUpgrade webhooks and dispatches upgrade_webhook_handler.sh.
|
||||
#
|
||||
# Runs outside Unraid nginx — no session auth required. The shared secret in
|
||||
# the webhook URL is the only gate. Arrs on this host call:
|
||||
#
|
||||
# http://<HOST_LAN_IP>:<WEBHOOK_PORT>/webhook?key=<WEBHOOK_SECRET>
|
||||
#
|
||||
# Runs as a continuous script started by array_started.sh. Execs node which
|
||||
# replaces this process — the PID stays the same for array_started.sh's check.
|
||||
#
|
||||
# Uses Node.js instead of php -S: php -S on Unraid PHP 8.4 silently drops
|
||||
# POST request bodies, making webhook payloads arrive empty.
|
||||
#
|
||||
# If WEBHOOK_SECRET is empty in master.conf: generates and saves one, then starts.
|
||||
# If WEBHOOK_PORT is 0: exits cleanly (disables the listener).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Check WEBHOOK_PORT — exit cleanly if 0 (listener disabled)
|
||||
# 2. acquire_lock "continuous" — exit cleanly if a healthy instance is already running
|
||||
# 3. Check WEBHOOK_SECRET — generate and persist one if empty
|
||||
# 4. exec node webhook_listener.js — replaces this process; PID stays the same
|
||||
#
|
||||
# exec is intentional: array_started.sh tracks the PID of this script to check
|
||||
# whether the listener is running. exec preserves that PID across the hand-off
|
||||
# to node.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Node.js Over php -S
|
||||
# php -S on Unraid PHP 8.4 silently drops POST request bodies — webhooks arrive
|
||||
# empty and the handler has no payload to act on. Node.js handles POST bodies
|
||||
# correctly and has no equivalent silent-drop behaviour.
|
||||
#
|
||||
# Runs Outside nginx
|
||||
# The listener binds directly to WEBHOOK_PORT — no nginx proxy, no session auth.
|
||||
# The shared secret in the URL query string is the only gate. This keeps the
|
||||
# webhook path independent of the auth stack.
|
||||
#
|
||||
# exec Preserves PID
|
||||
# The script execs into node rather than forking it. array_started.sh stores the
|
||||
# PID of this script to check liveness — exec ensures that PID continues to
|
||||
# refer to the running node process after the hand-off.
|
||||
#
|
||||
# Continuous-Mode Lock, Not Just PID Tracking
|
||||
# array_started.sh only checks whether ITS launch attempt is still alive after
|
||||
# 1s — it has no idea a previous instance might already be listening (e.g. an
|
||||
# array stop/start that didn't kill the old node process). Without its own
|
||||
# guard, a relaunch would exec straight into node, hit EADDRINUSE on the port,
|
||||
# exit 1 almost immediately, and array_started.sh would log a false failure
|
||||
# for a listener that was actually still healthy. acquire_lock "continuous"
|
||||
# detects the live instance and exits 0 instead.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Writes /var/log/varaverk and persists the generated secret into master.conf.
|
||||
#
|
||||
# WEBHOOK_PORT=0 Gate
|
||||
# Exits cleanly before any setup if the listener is disabled.
|
||||
#
|
||||
# node Presence Check
|
||||
# node is exec'd at the end of this script. Checking up front fails with a clear
|
||||
# reason at array start rather than an exec error buried in the log after setup.
|
||||
#
|
||||
# openssl Presence Check
|
||||
# Checked before attempting to generate a secret, on the path that needs it.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock "continuous" exits 0 cleanly if a healthy instance is already
|
||||
# listening, instead of relaunching into an EADDRINUSE port conflict. An array
|
||||
# stop/start without a full reboot leaves the old node process alive, and without
|
||||
# this array_started.sh would log a false failure.
|
||||
#
|
||||
# Secret Auto-Generation
|
||||
# WEBHOOK_SECRET generated via openssl rand if empty, and persisted to master.conf
|
||||
# so restarts reuse it.
|
||||
#
|
||||
# Secret Persistence Verification
|
||||
# The write-back is confirmed by re-reading master.conf. If it did not land, the
|
||||
# secret would exist only in this process and be regenerated on the next start,
|
||||
# silently invalidating the key already registered in the arrs — so this aborts
|
||||
# loudly rather than starting with a secret that will not survive a restart.
|
||||
#
|
||||
# Shared Secret Gate
|
||||
# The webhook URL must include ?key=<WEBHOOK_SECRET>; requests without a valid key
|
||||
# are rejected by the Node.js server.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# WEBHOOK_PORT — port the listener binds to; 0 = disabled
|
||||
# WEBHOOK_SECRET — shared secret for URL auth; auto-generated if empty
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# start_webhook_listener.sh
|
||||
# Started automatically at array start via ARRAY_START_SCRIPTS.
|
||||
# Exits immediately if WEBHOOK_PORT=0.
|
||||
#
|
||||
# To stop:
|
||||
# pkill -f webhook_listener.js
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
|
||||
source "$ECOSYSTEM_ROOT/load_config.sh"
|
||||
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "${WEBHOOK_PORT:-0}" -eq 0 ]] && {
|
||||
echo "[webhook] WEBHOOK_PORT=0 — listener disabled"
|
||||
exit 0
|
||||
}
|
||||
|
||||
# node is exec'd at the end of this script — checking here fails with a clear reason at
|
||||
# array start instead of an exec error buried in the log after all the setup has run.
|
||||
if ! command -v node >/dev/null 2>&1; then
|
||||
error "node not found — required to run webhook_listener.js"
|
||||
notify "Webhook listener failed to start on $(hostname) — node not installed" \
|
||||
"Webhook Listener" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Skip gracefully if a healthy instance is already listening — otherwise a
|
||||
# restart that doesn't kill the old node process (array stop/start without a
|
||||
# full reboot) hits EADDRINUSE and array_started.sh logs a false failure.
|
||||
acquire_lock "continuous"
|
||||
|
||||
# ── Auto-generate secret if not yet set ─────────────────────────────────────
|
||||
if [[ -z "${WEBHOOK_SECRET:-}" ]]; then
|
||||
if ! command -v openssl >/dev/null 2>&1; then
|
||||
error "openssl not found — cannot generate WEBHOOK_SECRET"
|
||||
notify "Webhook listener failed to start on $(hostname) — openssl not installed" \
|
||||
"Webhook Listener" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
GENERATED=$(openssl rand -hex 32)
|
||||
MASTER_CONF="$ECOSYSTEM_ROOT/Configurations/master.conf"
|
||||
sed -i "s/WEBHOOK_SECRET=\"\"/WEBHOOK_SECRET=\"$GENERATED\"/" "$MASTER_CONF"
|
||||
WEBHOOK_SECRET="$GENERATED"
|
||||
|
||||
# If the sed didn't match, the secret only exists in this process. The listener would
|
||||
# come up, then regenerate a different secret on the next start — silently invalidating
|
||||
# the key already registered in the arrs. Fail loudly instead.
|
||||
if ! grep -q "WEBHOOK_SECRET=\"$GENERATED\"" "$MASTER_CONF" 2>/dev/null; then
|
||||
error "Generated WEBHOOK_SECRET but could not persist it to $MASTER_CONF"
|
||||
error "Set WEBHOOK_SECRET manually — a non-persisted secret changes on every restart"
|
||||
notify "Webhook secret not persisted on $(hostname) — set WEBHOOK_SECRET manually" \
|
||||
"Webhook Listener" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "[webhook] Generated WEBHOOK_SECRET — run Tools/webhook_setup.sh to register in arrs"
|
||||
fi
|
||||
|
||||
mkdir -p /var/log/varaverk
|
||||
|
||||
exec node "$ECOSYSTEM_ROOT/Arrs_Stack/webhook_listener.js" \
|
||||
"$WEBHOOK_PORT" "$WEBHOOK_SECRET" "$ECOSYSTEM_ROOT" \
|
||||
>> /var/log/varaverk/upgrade_webhook.log 2>&1
|
||||
Executable
+258
@@ -0,0 +1,258 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================= Upgrade Webhook Handler ============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Triggered by Sonarr/Radarr/Lidarr OnUpgrade webhook (via webhook_listener.js).
|
||||
# Pushes the upgraded item folder to every other mesh node immediately, then
|
||||
# triggers a library rescan on each remote arr so it accepts the new file as
|
||||
# ground truth without initiating a redundant quality search.
|
||||
#
|
||||
# Closes the propagation window: without this, a remote node that already has
|
||||
# the 720p copy will see the 1080p tagged in arr_sync but not on disk and
|
||||
# begin searching — a search it will never win because we already have it.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Invoked per import by webhook_listener.js with <arr_type> <item_path>:
|
||||
#
|
||||
# 1. Validate
|
||||
# → arr_type must be sonarr|radarr|lidarr; item_path must exist and be a safe
|
||||
# absolute path
|
||||
#
|
||||
# 2. Resolve arr specifics
|
||||
# → API port, API version and rescan command for that arr type
|
||||
#
|
||||
# 3. Discover remote nodes
|
||||
# → discover_remote_nodes(); no remotes configured means exit cleanly
|
||||
#
|
||||
# 4. Per remote node, independently:
|
||||
# a. Resolve its Tailscale IP — unresolvable skips that node
|
||||
# b. rsync the single item to the same absolute path (--no-delete)
|
||||
# c. Skip the rescan if rsync failed — never scan a partial file
|
||||
# d. Trigger the arr's refresh command, cache-first API key with SSH fallback
|
||||
#
|
||||
# One failing node is counted and skipped; the rest still receive the upgrade.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Closes the Propagation Window
|
||||
# arr_sync runs every 4 hours. Without this handler, a remote node that already
|
||||
# has the old version sees the upgrade tagged in arr_sync but the new file not
|
||||
# yet on disk, and initiates a redundant quality search — a search it will never
|
||||
# win because this host already has the file. Immediate push eliminates that window.
|
||||
#
|
||||
# Rescan as Ground Truth
|
||||
# Pushing the file is not enough — the remote arr must also be told the file
|
||||
# exists. Triggering a rescan makes the remote accept the pushed file as the
|
||||
# current version without starting a new search.
|
||||
#
|
||||
# Cache-First API Key Lookup
|
||||
# Remote arr API keys are read from conf if cached; otherwise fetched via SSH
|
||||
# from the remote's config.xml. This avoids storing secrets redundantly while
|
||||
# keeping API calls fast on hosts where the key is already known.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# rsync runs over SSH as root and writes to root@remote at the same absolute path.
|
||||
#
|
||||
# Argument Validation
|
||||
# Exits with a usage message if arr_type or item_path is missing, and rejects an
|
||||
# arr_type outside sonarr|radarr|lidarr rather than defaulting to one.
|
||||
#
|
||||
# Path Existence
|
||||
# Exits if item_path is not a directory on disk.
|
||||
#
|
||||
# Item Path Depth Guard
|
||||
# item_path must be an absolute path at least three levels deep. It arrives from the
|
||||
# arr's webhook payload and is rsynced to the same path on the partner, so a truncated
|
||||
# or malformed value would push a system directory — or the filesystem root — onto the
|
||||
# remote. The existence check alone does not catch this, because / is a directory.
|
||||
#
|
||||
# No Lock — Deliberate
|
||||
# This is an event handler invoked per import by webhook_listener.js. Concurrent
|
||||
# upgrades are normal and expected. A default lock would silently drop overlapping
|
||||
# events, and a waiting lock would queue them behind a slow transfer, so neither is
|
||||
# used: each invocation rsyncs a different item path and they do not contend.
|
||||
#
|
||||
# Tailscale Resolution
|
||||
# Skips a node if its Tailscale IP cannot be resolved, rather than attempting the
|
||||
# transfer against an unresolved or stale address.
|
||||
#
|
||||
# rsync Exit Check
|
||||
# The rescan is only triggered if rsync succeeded. A failed transfer never causes the
|
||||
# remote arr to scan a partial file into its library.
|
||||
#
|
||||
# No Delete on Push
|
||||
# rsync runs with --no-delete. This pushes one upgraded item; it is not a mirror, and
|
||||
# must never remove content on the partner that this run does not know about.
|
||||
#
|
||||
# SSH Fallback
|
||||
# If no cached API key is available, falls back to SSH to read config.xml on the
|
||||
# remote rather than failing the rescan step.
|
||||
#
|
||||
# Per-Node Isolation
|
||||
# One unreachable or failing node is counted and skipped; the remaining nodes still
|
||||
# receive the upgrade.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST* — host names used to discover remote nodes
|
||||
# HOST*_SONARR_API_KEY — cached API key for direct HTTP rescan (optional)
|
||||
# HOST*_RADARR_API_KEY — cached API key for direct HTTP rescan (optional)
|
||||
# HOST*_LIDARR_API_KEY — cached API key for direct HTTP rescan (optional)
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# SSH_KEY — SSH key path for rsync and SSH fallback
|
||||
# ARR_SYNC_CONNECT_TIMEOUT — SSH connect timeout in seconds
|
||||
# DOCKER_APPDATA_BASE — base path for reading arr config.xml on remote
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# upgrade_webhook_handler.sh <arr_type> <item_path>
|
||||
#
|
||||
# arr_type — sonarr | radarr | lidarr
|
||||
# item_path — absolute path to the series/movie/artist folder on local disk
|
||||
# (series.path from Sonarr, movie.folderPath from Radarr,
|
||||
# artist.path from Lidarr)
|
||||
#
|
||||
# Called by webhook_listener.js — not intended for direct invocation outside testing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
ARR_TYPE="${1:-}"
|
||||
ITEM_PATH="${2:-}"
|
||||
|
||||
[[ -z "$ARR_TYPE" || -z "$ITEM_PATH" ]] && {
|
||||
echo "Usage: upgrade_webhook_handler.sh <arr_type> <item_path>" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
[[ -d "$ITEM_PATH" ]] || { echo "Path not found: $ITEM_PATH" >&2; exit 1; }
|
||||
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
echo "Must be run as root" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ITEM_PATH is rsynced to root@remote at the same absolute path. It arrives from the arr's
|
||||
# webhook payload, so a malformed or truncated value would push a system directory — or the
|
||||
# filesystem root — onto the partner. -d alone does not catch that: / is a directory.
|
||||
_depth="${ITEM_PATH//[^\/]/}"
|
||||
if [[ "$ITEM_PATH" != /* || "${#_depth}" -lt 3 ]]; then
|
||||
echo "Refusing unsafe item path: '$ITEM_PATH' — expected an absolute path at least 3 levels deep" >&2
|
||||
exit 1
|
||||
fi
|
||||
unset _depth
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
detect_hosts
|
||||
|
||||
# ── Arr type → API port / version / rescan command ───────────────────────────────────────────
|
||||
case "$ARR_TYPE" in
|
||||
sonarr) PORT=8989; API_VER="v3"; RESCAN_CMD="RefreshSeries" ;;
|
||||
radarr) PORT=7878; API_VER="v3"; RESCAN_CMD="RefreshMovie" ;;
|
||||
lidarr) PORT=8686; API_VER="v1"; RESCAN_CMD="RefreshArtist" ;;
|
||||
*) echo "Unknown arr type: $ARR_TYPE" >&2; exit 1 ;;
|
||||
esac
|
||||
|
||||
# ── Remote node list ──────────────────────────────────────────────────────────────────────────
|
||||
discover_remote_nodes
|
||||
|
||||
if [[ "${#REMOTE_NODES[@]}" -eq 0 ]]; then
|
||||
echo "No remote nodes configured — nothing to push"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
ENCODED=$(printf '{"name":"%s"}' "$RESCAN_CMD" | base64 -w0)
|
||||
ITEM_NAME=$(basename "$ITEM_PATH")
|
||||
|
||||
echo "[$(date '+%H:%M:%S')] Upgrade push: ${ARR_TYPE} — ${ITEM_NAME}"
|
||||
echo " Path: $ITEM_PATH"
|
||||
echo " Targets: ${REMOTE_NODES[*]}"
|
||||
|
||||
NODE_FAIL=0
|
||||
|
||||
# ── Push and rescan each remote ───────────────────────────────────────────────────────────────
|
||||
for node_id in "${REMOTE_NODES[@]}"; do
|
||||
node_name="${!node_id}"
|
||||
node_ip=$(resolve_tailscale_ip "$node_name") || {
|
||||
echo " [${node_name}] Cannot resolve Tailscale IP — skipping"
|
||||
(( NODE_FAIL++ ))
|
||||
continue
|
||||
}
|
||||
|
||||
# ── Targeted rsync — push only this item, no delete ──────────────────────────────────────
|
||||
echo " [${node_name}] rsync ${ITEM_NAME}..."
|
||||
rsync_out=$(rsync -av --no-delete \
|
||||
-e "ssh -i $SSH_KEY -T -o Compression=no -o IPQoS=throughput -o StrictHostKeyChecking=no" \
|
||||
"${ITEM_PATH}/" \
|
||||
"root@${node_ip}:${ITEM_PATH}/" 2>&1)
|
||||
rsync_exit=$?
|
||||
transferred=$(echo "$rsync_out" | awk '/Total transferred file size:/{gsub(/,/,"",$NF); gsub(/[^0-9]/,"",$NF); print $NF+0}')
|
||||
|
||||
if [[ "$rsync_exit" -ne 0 ]]; then
|
||||
echo " [${node_name}] rsync failed (exit ${rsync_exit}) — skipping rescan"
|
||||
(( NODE_FAIL++ ))
|
||||
continue
|
||||
fi
|
||||
echo " [${node_name}] rsync done (${transferred:-0} bytes)"
|
||||
|
||||
# ── Trigger arr rescan on remote — cache-first, SSH fallback ─────────────────────────────
|
||||
_kvar="${node_id}_${ARR_TYPE^^}_API_KEY"
|
||||
cached_key="${!_kvar:-}"
|
||||
|
||||
if [[ -n "$cached_key" ]]; then
|
||||
body=$(printf '%s' "$ENCODED" | base64 -d)
|
||||
http_code=$(curl -sf -o /dev/null -w '%{http_code}' -X POST \
|
||||
-H "X-Api-Key: $cached_key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body" \
|
||||
"http://${node_ip}:${PORT}/api/${API_VER}/command" 2>/dev/null)
|
||||
else
|
||||
config_xml="${DOCKER_APPDATA_BASE}/${ARR_TYPE^}/config.xml"
|
||||
http_code=$(ssh -i "$SSH_KEY" -o ConnectTimeout="$ARR_SYNC_CONNECT_TIMEOUT" \
|
||||
root@"$node_ip" bash <<REMOTE 2>/dev/null
|
||||
KEY=\$(grep -oP '(?<=<ApiKey>)[^<]+' '${config_xml}' 2>/dev/null)
|
||||
[[ -z "\$KEY" ]] && exit 1
|
||||
BODY=\$(printf '%s' '${ENCODED}' | base64 -d)
|
||||
curl -sf -o /dev/null -w '%{http_code}' -X POST \
|
||||
-H "X-Api-Key: \$KEY" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "\$BODY" \
|
||||
"http://localhost:${PORT}/api/${API_VER}/command"
|
||||
REMOTE
|
||||
)
|
||||
fi
|
||||
|
||||
if [[ "$http_code" == "201" || "$http_code" == "200" ]]; then
|
||||
echo " [${node_name}] ${RESCAN_CMD} triggered ✅"
|
||||
else
|
||||
echo " [${node_name}] ${RESCAN_CMD} failed (HTTP ${http_code:-timeout})"
|
||||
(( NODE_FAIL++ ))
|
||||
fi
|
||||
done
|
||||
|
||||
echo "[$(date '+%H:%M:%S')] Done — ${ITEM_NAME}"
|
||||
|
||||
[[ "$NODE_FAIL" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
@@ -0,0 +1,108 @@
|
||||
#!/usr/bin/env node
|
||||
// Standalone arr upgrade webhook listener.
|
||||
// Started by start_webhook_listener.sh via `node webhook_listener.js`.
|
||||
// Lives outside Unraid nginx — no session auth. Secret in URL is the only gate.
|
||||
//
|
||||
// Usage: node webhook_listener.js <port> <secret> <scripts_dir>
|
||||
|
||||
'use strict';
|
||||
|
||||
const http = require('http');
|
||||
const { exec } = require('child_process');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
|
||||
const [,, port, secret, scriptsDir] = process.argv;
|
||||
|
||||
if (!port || !secret || !scriptsDir) {
|
||||
process.stderr.write('Usage: webhook_listener.js <port> <secret> <scripts_dir>\n');
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
const LOG_FILE = '/var/log/varaverk/upgrade_webhook.log';
|
||||
const HANDLER = path.join(scriptsDir, 'Arrs_Stack', 'upgrade_webhook_handler.sh');
|
||||
|
||||
function log(msg) {
|
||||
const ts = new Date().toTimeString().slice(0, 8);
|
||||
const line = `[${ts}] ${msg}\n`;
|
||||
fs.appendFile(LOG_FILE, line, () => {});
|
||||
}
|
||||
|
||||
function send(res, code, obj) {
|
||||
const body = JSON.stringify(obj);
|
||||
res.writeHead(code, { 'Content-Type': 'application/json', 'Content-Length': Buffer.byteLength(body) });
|
||||
res.end(body);
|
||||
}
|
||||
|
||||
const server = http.createServer((req, res) => {
|
||||
const url = new URL(req.url, `http://localhost`);
|
||||
const key = url.searchParams.get('key');
|
||||
|
||||
if (key !== secret) {
|
||||
send(res, 403, { ok: false, error: 'Forbidden' });
|
||||
return;
|
||||
}
|
||||
|
||||
if (req.method !== 'POST') {
|
||||
send(res, 405, { ok: false, error: 'POST only' });
|
||||
return;
|
||||
}
|
||||
|
||||
let body = '';
|
||||
req.on('data', chunk => { body += chunk; });
|
||||
req.on('end', () => {
|
||||
let payload;
|
||||
try { payload = JSON.parse(body); }
|
||||
catch (_) { send(res, 400, { ok: false, error: 'Invalid JSON' }); return; }
|
||||
|
||||
const event = payload.eventType ?? '';
|
||||
const isUpgrade = payload.isUpgrade ?? false;
|
||||
|
||||
if (event === 'Test') {
|
||||
send(res, 200, { ok: true, message: 'Webhook connected — upgrade propagation active' });
|
||||
return;
|
||||
}
|
||||
|
||||
if (event !== 'Download' || !isUpgrade) {
|
||||
send(res, 200, { ok: true, skipped: event });
|
||||
return;
|
||||
}
|
||||
|
||||
let arrType, itemPath;
|
||||
if (payload.series?.path) {
|
||||
arrType = 'sonarr';
|
||||
itemPath = payload.series.path;
|
||||
} else if (payload.movie?.folderPath) {
|
||||
arrType = 'radarr';
|
||||
itemPath = payload.movie.folderPath;
|
||||
} else if (payload.artist?.path) {
|
||||
arrType = 'lidarr';
|
||||
itemPath = payload.artist.path;
|
||||
} else {
|
||||
send(res, 400, { ok: false, error: 'Unrecognised payload structure' });
|
||||
return;
|
||||
}
|
||||
|
||||
if (!itemPath || !itemPath.startsWith('/') || itemPath.includes('..') || /[\x00\n\r]/.test(itemPath)) {
|
||||
send(res, 400, { ok: false, error: 'Unsafe path' });
|
||||
return;
|
||||
}
|
||||
|
||||
log(`Upgrade: ${arrType} — ${path.basename(itemPath)}`);
|
||||
send(res, 200, { ok: true, arr: arrType, path: itemPath });
|
||||
|
||||
const cmd = `bash ${JSON.stringify(HANDLER)} ${JSON.stringify(arrType)} ${JSON.stringify(itemPath)}`;
|
||||
const out = fs.createWriteStream(LOG_FILE, { flags: 'a' });
|
||||
exec(cmd, { stdio: ['ignore', out, out] });
|
||||
});
|
||||
});
|
||||
|
||||
server.listen(parseInt(port, 10), '0.0.0.0', () => {
|
||||
log(`Webhook listener started on port ${port}`);
|
||||
process.stdout.write(`[webhook] Listening on port ${port}\n`);
|
||||
});
|
||||
|
||||
server.on('error', err => {
|
||||
process.stderr.write(`[webhook] Server error: ${err.message}\n`);
|
||||
process.exit(1);
|
||||
});
|
||||
@@ -0,0 +1,15 @@
|
||||
# Configurations/
|
||||
|
||||
Runtime home for the live conf files. These are gitignored — never committed.
|
||||
|
||||
| File | Source |
|
||||
|------|--------|
|
||||
| `master.conf` | Seeded from `Deployment/master.conf.template` on first install |
|
||||
| `host1.conf` | Seeded from `Deployment/host.conf.template` on first install |
|
||||
| `host2.conf` | Seeded from `Deployment/host.conf.template` on first install |
|
||||
|
||||
**Templates and tooling live in `Deployment/`:**
|
||||
- `Deployment/host.conf.template` — canonical template for new host confs
|
||||
- `Deployment/master.conf.template` — canonical template for master.conf
|
||||
- `Deployment/conf_upgrade.sh` — merges schema changes into existing confs
|
||||
- `Deployment/conf_populate.sh` — auto-detects and fills arr keys on first run
|
||||
@@ -1,769 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================== HOST1 CONFIGURATION — unRAID-Gmer4Lfe ============================
|
||||
# ==============================================================================================
|
||||
# HOST1-specific variables — credentials, container names, share paths, failover lists.
|
||||
# Sourced after master.conf — values here extend shared profile arrays and add HOST1-specific
|
||||
# identity, credentials, and container configuration.
|
||||
#
|
||||
# Sparse checkout (git) ensures HOST2 never receives this file.
|
||||
# HOST2 never sees HOST1 credentials — clean separation at the file level.
|
||||
#
|
||||
# DO NOT put shared config here — thresholds, toggles, profiles belong in master.conf.
|
||||
# DO NOT put HOST2 variables here — they belong in host2.conf.
|
||||
#
|
||||
# ── INDEX ─────────────────────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# ── IDENTITY & CONNECTIVITY ────────────────────────────────────────────────────────────────
|
||||
# IDENTITY hostname, SSH key
|
||||
# EMBY container name, URL, API key
|
||||
# NOTIFICATIONS Discord webhook
|
||||
# PARTNERSHIP auth containers, backup paths
|
||||
#
|
||||
# ── RSYNC ──────────────────────────────────────────────────────────────────────────────────
|
||||
# DAILY SYNC SHARES media shares HOST1 owns and pushes to HOST2
|
||||
# WEEKLY SYNC SHARES appdata shares synced weekly (Sunday 2:30am)
|
||||
# CRITICAL SYNC SHARES appdata shares synced every 30 minutes
|
||||
# BACKUP VERIFY shares for checksum verification against remote
|
||||
# HOST1 RSYNC PROFILE host1-appdata profile for HOST1-specific appdata syncs
|
||||
#
|
||||
# ── DOCKER ─────────────────────────────────────────────────────────────────────────────────
|
||||
# DOCKER DAILY RESTART containers restarted daily
|
||||
# DOCKER WEEKLY RESTART containers restarted weekly
|
||||
# DOCKER WATCHDOG memory limits, health URLs, required containers, ignore list
|
||||
# DOCKER NETWORK CONNECT networks and containers for docker_network_connect.sh
|
||||
#
|
||||
# ── FALLBACK ───────────────────────────────────────────────────────────────────────────────
|
||||
# DDNS DDNS containers managed by HOST1
|
||||
# INTERNET LOSS containers stopped when internet is lost
|
||||
# FALLBACK TIERS what HOST1 runs for HOST2 per tier
|
||||
# TIER DELAYS how long HOST1 must be down before each tier activates on HOST2
|
||||
# RSYNC WRITEBACK HOST1 appdata synced back on handback
|
||||
#
|
||||
# ── MEDIA ──────────────────────────────────────────────────────────────────────────────────
|
||||
# MEDIA PERMISSIONS share list for media_shares_permissions.sh
|
||||
# MEDIA CLEANER folder lists for media_cleaner.sh
|
||||
#
|
||||
# ── MONITORS ───────────────────────────────────────────────────────────────────────────────
|
||||
# CERTIFICATE MONITOR domains checked for SSL expiry
|
||||
# SMART HEALTH drives to skip in SMART monitoring
|
||||
# ZFS REPORT pools to exclude from ZFS health report
|
||||
#
|
||||
# ── TRANSCODES ─────────────────────────────────────────────────────────────────────────────
|
||||
# TRANSCODES ramdisk size, thresholds, SSD path, server array
|
||||
#
|
||||
# ── ARR STACK ──────────────────────────────────────────────────────────────────────────────
|
||||
# DOWNLOADERS slskd, SABnzbd, qBittorrent credentials and URLs
|
||||
# LIDARR URL, API key, path map
|
||||
# SONARR URL, API key, path map
|
||||
# RADARR URL, API key, path map
|
||||
# ARR RECOVERY per-arr recovery toggles
|
||||
#
|
||||
# ── SYSTEM WATCHDOG ────────────────────────────────────────────────────────────────────────
|
||||
# SYSTEM WATCHDOG per-host check toggles and NIC configuration
|
||||
#
|
||||
# ── RESOURCE MANAGER ───────────────────────────────────────────────────────────────────────
|
||||
# RESOURCE MANAGER containers paused/stopped under memory pressure
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
# ==============================================================================================
|
||||
# ── IDENTITY & CONNECTIVITY ───────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Identity ━━━
|
||||
# HOST1 hostname lives in master.conf (not a credential — safe for all servers).
|
||||
# SSH key used for all server-to-server operations — rsync, failover container commands.
|
||||
# Must be in /root/.ssh/ and authorised in HOST2's /root/.ssh/authorized_keys.
|
||||
HOST1_SSH_KEY="/root/.ssh/gmer4lfe_rsync_automation"
|
||||
HOST1_OWNER="gmer4lfe"
|
||||
HOST1_OWNER_EMAIL="gmer4lfe@gmail.com"
|
||||
|
||||
# ━━━ Emby ━━━
|
||||
# Referenced by transcode_manager.sh, emby_session_report.sh, emby_database_repair.sh,
|
||||
# weekly_sync_maintenance.sh, and HOST1_TRANSCODE_SERVERS below.
|
||||
# API key: Emby Dashboard → API Keys → + New Key
|
||||
HOST1_EMBY_CONTAINER="Emby"
|
||||
HOST1_EMBY_URL="http://localhost:8096"
|
||||
HOST1_EMBY_API_KEY="0c27448d93a7431f9ac63569f7655829"
|
||||
|
||||
# ━━━ Jellyfin ━━━
|
||||
# API key: Jellyfin Dashboard → Administration → API Keys → + New Key
|
||||
HOST1_JELLYFIN_CONTAINER="Jellyfin"
|
||||
HOST1_JELLYFIN_URL="http://localhost:8095"
|
||||
HOST1_JELLYFIN_API_KEY="956d0168987f4e4680626653abb080f0"
|
||||
|
||||
# ━━━ Gitea ━━━
|
||||
# Personal access token for gitea_ssh_setup.sh — registers this server's SSH public key
|
||||
# with Gitea so git operations use key auth instead of passwords.
|
||||
# Create in Gitea: Settings → Applications → Generate Token → scope: write:user
|
||||
HOST1_GITEA_API_TOKEN=""
|
||||
|
||||
# ━━━ Notifications ━━━
|
||||
# Discord webhook — leave blank to disable.
|
||||
# Per-host so HOST1 and HOST2 can post to different channels or only one server notifies.
|
||||
HOST1_DISCORD_WEBHOOK=""
|
||||
|
||||
# ━━━ Partnership ━━━
|
||||
# HOST1 is always the owner (source of truth) unless --transfer has been run.
|
||||
# See README-Partnership.md and master.conf PARTNERSHIP section for full lifecycle docs.
|
||||
|
||||
# Auth containers reconfigured on onboard/offboard.
|
||||
# Format: "ContainerName|WebUIPort"
|
||||
# On onboard → WebUI pointed at owner's Tailscale IP (mirror clicks NPM, gets owner's auth)
|
||||
# On offboard → WebUI pointed back at localhost
|
||||
HOST1_PARTNERSHIP_AUTH_WEBUIS=(
|
||||
"NginxProxyManager|81"
|
||||
"Lldap-Gmer4Lfe|17170"
|
||||
"Authelia|9091"
|
||||
"Authelia-Secondary|9092"
|
||||
)
|
||||
|
||||
# XML templates (from this server's templates-user/) pushed to mirror during onboard.
|
||||
# These become the mirror's active auth stack, backed by the rsync-synced appdata.
|
||||
# Update filename if Lldap is renamed to drop the host suffix.
|
||||
HOST1_PARTNERSHIP_AUTH_STACK=(
|
||||
# Dependencies first — Mariadb/Redis must be healthy before Authelia starts
|
||||
"my-Mariadb-Authelia.xml"
|
||||
"my-Mariadb-Authelia-Secondary.xml"
|
||||
"my-Redis-Authelia.xml"
|
||||
"my-Redis-Authelia-Secondary.xml"
|
||||
# Auth apps — deployed after their deps are confirmed healthy
|
||||
"my-Authelia.xml"
|
||||
"my-Authelia-Secondary.xml"
|
||||
"my-NginxProxyManager.xml"
|
||||
"my-Lldap-Gmer4Lfe.xml"
|
||||
# Source of truth — must be available on HOST2 independently of the auth stack
|
||||
"my-Gitea.xml"
|
||||
)
|
||||
|
||||
# XML templates pushed to mirror for the arr stack during onboard.
|
||||
# Deps (e.g. databases) first if any — same ordering rule as auth stack.
|
||||
HOST1_PARTNERSHIP_ARR_STACK=(
|
||||
# "my-Sonarr.xml"
|
||||
# "my-Radarr.xml"
|
||||
# "my-Lidarr.xml"
|
||||
# "my-Prowlarr.xml"
|
||||
# "my-Bazarr.xml"
|
||||
)
|
||||
|
||||
# Paths HOST2 should collect during the grace window after offboard.
|
||||
# Notified on offboard — no auto-deletion, HOST2 must collect manually within PARTNERSHIP_GRACE_HOURS.
|
||||
HOST1_PARTNERSHIP_MIRROR_BACKUPS=(
|
||||
# "/mnt/user/appdata-Fallback/Jayred365-Emby"
|
||||
)
|
||||
|
||||
# Containers parked on this server when partnership is active.
|
||||
# Stopped on onboard (owner deploys its stack instead), restarted on offboard.
|
||||
HOST1_PARTNERSHIP_OWN_CONTAINERS=(
|
||||
# "Emby"
|
||||
# "NginxProxyManager"
|
||||
)
|
||||
|
||||
# Emby admin provisioning — toggle is owner-only, credentials are per-host.
|
||||
# Owner enables/disables the feature. Each host sets the account they want on the shared Emby.
|
||||
# On onboard: owner reads mirror's HOST*_PARTNERSHIP_EMBY_ADMIN_* and creates that account.
|
||||
# On offboard: account is deleted. Username collision → onboard exits with error.
|
||||
HOST1_PARTNERSHIP_PROVISION_EMBY_ADMIN=false # owner controls whether Emby is shared
|
||||
HOST1_PARTNERSHIP_EMBY_PORT=8096
|
||||
HOST1_PARTNERSHIP_EMBY_ADMIN_USER="" # this server's desired Emby username
|
||||
HOST1_PARTNERSHIP_EMBY_ADMIN_PASS="" # this server's desired Emby password
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RSYNC ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Daily Sync Shares ━━━
|
||||
# Shares HOST1 pushes to all other nodes every night (1am via daily_sync_maintenance.sh).
|
||||
# Mesh model: every node pushes every media share — no ownership, no mirrors.
|
||||
# arr_sync ensures all arr libraries converge (union). rsync spreads files (additive, no --delete).
|
||||
# arr_cleanup removes true orphans based on local arr state.
|
||||
# Any node can download content to any share — it propagates to all nodes on the next cycle.
|
||||
# Nextcloud is intentionally one-directional (HOST1→HOST2 offsite backup — not arr-managed).
|
||||
# Uses DEFAULT_RSYNC_OPTS from master.conf — no profile needed.
|
||||
# For shares needing container stops or custom options — add a profile in master.conf.
|
||||
HOST1_DAILY_SYNC_SHARES=(
|
||||
/mnt/user/Books
|
||||
/mnt/user/Intros
|
||||
/mnt/user/Kids_Movies
|
||||
/mnt/user/Kids_Tv_Shows
|
||||
/mnt/user/Movies
|
||||
/mnt/user/Music
|
||||
/mnt/user/Music_Videos
|
||||
/mnt/user/Nextcloud
|
||||
/mnt/user/stand-up_comedy
|
||||
/mnt/user/Sports
|
||||
/mnt/user/Tv_Shows
|
||||
/mnt/user/Anime_Shows-Old
|
||||
/mnt/user/Anime_Movies-Old
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Shows
|
||||
)
|
||||
|
||||
# Personal encrypted shares — synced for offsite backup, independent of media shares.
|
||||
# ZFS encrypted at dataset level — remote receives encrypted blocks, cannot read content.
|
||||
# See README-Rsync_Setup.md for ZFS encryption setup before uncommenting.
|
||||
HOST1_PERSONAL_SHARES=(
|
||||
# /mnt/user/HOST1-Personal # uncomment after creating encrypted dataset
|
||||
)
|
||||
|
||||
# ━━━ Weekly Sync Shares ━━━
|
||||
# Appdata shares synced during the weekly maintenance window (Sunday 2:30am).
|
||||
# Containers stopped both sides before sync — full clean state guaranteed.
|
||||
# Profiles drive container stops, excludes, and options — configured in master.conf RSYNC section.
|
||||
# Order matters — Emby first (larger transfer), then Critical-Data (auth stack).
|
||||
HOST1_WEEKLY_SYNC_SHARES=(
|
||||
"/mnt/user/Media_Server/Emby" # emby profile — full clean mirror
|
||||
"/mnt/user/appdata-Fallback/Critical-Data" # critical-data profile — auth stack
|
||||
)
|
||||
|
||||
# ━━━ Intermediate Sync Shares ━━━
|
||||
# Shares synced every 4 hours by intermediate_sync_maintenance.sh.
|
||||
# Uses DEFAULT_RSYNC_OPTS (no --delete) — for sub-daily propagation of metadata or watch state.
|
||||
# Full media share sync stays in the daily window. Leave empty to skip mid-day rsync entirely.
|
||||
HOST1_INTERMEDIATE_SYNC_SHARES=(
|
||||
# Add shares here to enable mid-day rsync
|
||||
# Example: "/mnt/user/Emby_Metadata"
|
||||
)
|
||||
|
||||
# ━━━ Critical Sync Shares ━━━
|
||||
# Appdata shares synced every 30 minutes by critical_sync_maintenance.sh.
|
||||
# Format: "/path/to/share" or "/path/to/share|profile-name"
|
||||
# Order matters — Critical-Data first (auth stack), then Emby dirty sync.
|
||||
HOST1_CRITICAL_SYNC_SHARES=(
|
||||
"/mnt/user/appdata-Fallback/Critical-Data|critical-fallback" # auth dirty sync — stays running
|
||||
"/mnt/user/Media_Server/Emby|emby-fallback" # Emby dirty sync — stays running
|
||||
)
|
||||
|
||||
# ━━━ Backup Verify ━━━
|
||||
# Shares verified by backup_verify.sh — random file checksum comparison against remote.
|
||||
# Leave empty to use HOST1_DAILY_SYNC_SHARES automatically.
|
||||
# Sample size and minimum file size defined in master.conf.
|
||||
HOST1_BACKUP_VERIFY_SHARES=(
|
||||
# leave empty to use HOST1_DAILY_SYNC_SHARES automatically
|
||||
)
|
||||
|
||||
# ━━━ HOST1 Rsync Profile — host1-appdata ━━━
|
||||
# HOST1-specific appdata sync profile — extends the shared PROFILE_* arrays in master.conf.
|
||||
# Use for appdata unique to HOST1 (Organizrv2, VaultWarden, UptimeKuma etc.)
|
||||
# Shared appdata (auth stack, Emby) use dedicated profiles defined in master.conf.
|
||||
# Run manually: bash Rsync/rsync.sh /mnt/user/appdata-Fallback/HOST1-Appdata --profile=host1-appdata
|
||||
PROFILE_RSYNC_OPTS[host1-appdata]="-av --info=progress2 --bwlimit=${PROFILE_BW_LIMIT[host1-appdata]:-8000}"
|
||||
PROFILE_BW_LIMIT[host1-appdata]=8000
|
||||
PROFILE_RETRY_COUNT[host1-appdata]=3
|
||||
PROFILE_SLEEP[host1-appdata]=300
|
||||
PROFILE_CRITICAL_CONTAINER_NAMES[host1-appdata]="Organizrv2-Gmer4Lfe UptimeKuma-Gmer4Lfe VaultWarden-Gmer4Lfe"
|
||||
PROFILE_DELAYED_CONTAINERS[host1-appdata]=""
|
||||
PROFILE_CONTAINER_DELAY[host1-appdata]=5
|
||||
PROFILE_EXCLUDE_DIRS[host1-appdata]="logs *.tmp"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── DOCKER ────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Docker Daily Restart ━━━
|
||||
# Containers restarted every day via DAILY_MAINTENANCE_SCRIPTS.
|
||||
# Dispatcharr degrades over time without restart — daily is intentional, not just housekeeping.
|
||||
# Order matters — auth stack first, then media services.
|
||||
HOST1_DAILY_RESTART_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
"Lldap-Gmer4Lfe"
|
||||
"Authelia"
|
||||
"Authelia-Secondary"
|
||||
"Dispatcharr-Iptv-Users"
|
||||
"Dispatcharr" # Live TV scheduler — degrades without daily restart
|
||||
"Dispatcharr-Basic"
|
||||
"ErsatzTV-Emby"
|
||||
"Slskd" # Soulseek connection drops after extended uptime; restart refreshes share index
|
||||
)
|
||||
|
||||
# ━━━ Docker Weekly Restart ━━━
|
||||
# Less critical services restarted weekly via WEEKLY_MAINTENANCE_SCRIPTS (Sunday 2:30am).
|
||||
# Containers already stopped for weekly sync — restart adds zero extra downtime.
|
||||
HOST1_WEEKLY_RESTART_CONTAINERS=(
|
||||
"NextCloud"
|
||||
"Organizrv2-Gmer4Lfe"
|
||||
"AdGuard-Home"
|
||||
"Immich-Gmer4Lfe"
|
||||
)
|
||||
|
||||
# ━━━ Docker Watchdog ━━━
|
||||
# Per-HOST1 container configuration for docker_watchdog.sh.
|
||||
# Shared thresholds and toggles live in master.conf.
|
||||
|
||||
# Memory hard limits in MB — immediate restart if exceeded.
|
||||
# Set at "container is clearly broken" not "container is busy".
|
||||
# 20GB=20480 18GB=18432 16GB=16384 12GB=12288 8GB=8192 4GB=4096 2GB=2048 1GB=1024
|
||||
declare -A HOST1_WATCHDOG_CONTAINERS=(
|
||||
["Emby"]=18432 # 18GB — large library + active transcodes
|
||||
["LidaTube"]=6144 # 6GB — memory leak over time
|
||||
["Tdarr"]=6144 # 6GB — encoding is memory intensive
|
||||
["Code-Server"]=1024 # 1GB — should never need more
|
||||
)
|
||||
|
||||
# HTTP health check URLs — checked every cycle, strike system before restart.
|
||||
# Only add containers with a meaningful web interface to check.
|
||||
declare -A HOST1_WATCHDOG_CONTAINER_URLS=(
|
||||
["Emby"]="http://localhost:8096"
|
||||
)
|
||||
|
||||
# Required containers — must always be running on HOST1.
|
||||
# Strike system before restart — repeated failures go on skip list, auto-clears on recovery.
|
||||
# Listed in dependency order — dependencies before dependents.
|
||||
HOST1_WATCHDOG_REQUIRED_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
"Lldap-Gmer4Lfe"
|
||||
"Mariadb-Authelia"
|
||||
"Mariadb-Authelia-Secondary"
|
||||
"Redis-Authelia"
|
||||
"Redis-Authelia-Secondary"
|
||||
"Authelia"
|
||||
"Authelia-Secondary"
|
||||
)
|
||||
|
||||
# Containers to skip in Tier 2 global scan — legitimately stopped or frequently restarting.
|
||||
# Watchdog leaves these alone entirely — no restart attempts, no crash loop tracking.
|
||||
HOST1_WATCHDOG_SCAN_IGNORE=(
|
||||
"DashGate"
|
||||
"PIA-WG-Config-Generator"
|
||||
"Aperture"
|
||||
"Aperture-Kids"
|
||||
"pgvector-18-Apeture-Kids"
|
||||
"Pgvector18-Aperture"
|
||||
)
|
||||
|
||||
# Dependency ordering — skip restarting a container if its dependency is also down.
|
||||
# Prevents watchdog from restarting Authelia before Mariadb is back up.
|
||||
# SPACE-SEPARATED STRINGS — converted to array at runtime.
|
||||
declare -A HOST1_WATCHDOG_DEPENDENCIES=(
|
||||
["Authelia"]="Mariadb-Authelia Redis-Authelia"
|
||||
["Authelia-Secondary"]="Mariadb-Authelia Redis-Authelia-Secondary"
|
||||
["NextCloud"]="Postgres-NextCloud"
|
||||
)
|
||||
|
||||
# Per-container appdata growth suppress ceilings in MB.
|
||||
# ONLY needed in specific cases — growth rate detection covers all containers automatically.
|
||||
# Use this when a container legitimately has large stable data and you want to guarantee
|
||||
# it never triggers a false-positive growth alert. Growth warnings are suppressed while the
|
||||
# container's dir stays below this ceiling; above it, warnings resume as normal.
|
||||
# 50GB=51200 25GB=25600 20GB=20480 15GB=15360 10GB=10240 5GB=5120
|
||||
declare -A HOST1_WATCHDOG_APPDATA_SIZES=(
|
||||
["Tdarr"]="25600" # 25GB — transcode cache grows legitimately during active jobs
|
||||
["7dtd"]="20480" # 20GB — game server world data, expected to be large
|
||||
)
|
||||
|
||||
# ━━━ Network Watchdog ━━━
|
||||
# Host-specific connectivity config for Watchdogs/System/network_watchdog.sh.
|
||||
HOST1_NETWORK_WATCHDOG_DDNS_DOMAIN="gmer4lfe.com"
|
||||
HOST1_NETWORK_WATCHDOG_DDNS_CONTAINER="Gmer4Lfe.com"
|
||||
HOST1_NETWORK_WATCHDOG_NPM_URL="https://gmer4lfe.com"
|
||||
|
||||
# ━━━ Docker Network Connect ━━━
|
||||
# Containers connected to custom networks at array start by docker_network_connect.sh.
|
||||
# Networks created if they don't exist — idempotent, safe to re-run.
|
||||
HOST1_NETWORK_CONNECT_CONTAINERS=(
|
||||
"memcached"
|
||||
"Npm-CrowdSec"
|
||||
)
|
||||
|
||||
HOST1_NETWORK_CONNECT_NETWORKS=(
|
||||
"high-availability"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FALLBACK ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ DDNS ━━━
|
||||
# DDNS containers HOST1 manages — started/stopped by fallback.sh per DDNS absolute rules:
|
||||
# Internet loss → stop immediately
|
||||
# Failover → HOST2 starts HOST1's DDNS as Tier 1 (before any other containers)
|
||||
# Handback → stop HOST1's DDNS on HOST2 → rsync → start containers → start local DDNS last
|
||||
HOST1_DDNS_CONTAINERS=(
|
||||
"Gmer4Lfe.com"
|
||||
)
|
||||
|
||||
# ━━━ Internet Loss ━━━
|
||||
# Containers stopped immediately on HOST1 when internet connection is lost.
|
||||
# Prevents external-facing services from operating without connectivity.
|
||||
FALLBACK_HOST1_STOP_ON_NO_NET=(
|
||||
"Gmer4Lfe.com"
|
||||
)
|
||||
|
||||
# ━━━ Fallback Tiers — HOST1 Runs for HOST2 ━━━
|
||||
# Containers HOST1 starts when HOST2 goes down.
|
||||
# Tier 1 is always immediate — vital services cannot wait.
|
||||
# Higher tiers activate after HOST2_TIER*_DELAY minutes (set in host2.conf).
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER1=(
|
||||
"Gmer4Lfe.us"
|
||||
"VaultWarden-Jayred365"
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER2=(
|
||||
# "container-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER3=(
|
||||
# "container-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER4=(
|
||||
# "container-placeholder"
|
||||
)
|
||||
|
||||
# ━━━ Tier Delays — HOST1's Containers on HOST2 ━━━
|
||||
# How long HOST1 must be down before each tier activates on HOST2 — in minutes.
|
||||
# Tier 1 is always immediate — no delay var needed.
|
||||
HOST1_TIER2_DELAY=240 # 4 hours — NextCloud, Immich
|
||||
HOST1_TIER3_DELAY=720 # 12 hours — secondary services
|
||||
HOST1_TIER4_DELAY=1440 # 24 hours — arrs + downloaders
|
||||
|
||||
# ━━━ Rsync Writeback — HOST1 Appdata Back on Handback ━━━
|
||||
# Syncs HOST1 appdata BACK to HOST1 when it comes back online after a failover.
|
||||
# Containers stopped before writeback — clean source, no competing writes.
|
||||
#
|
||||
# HOST1_TIER1_WRITEBACK_DELAY: short outages skip Tier 1 writeback — primary state
|
||||
# is more reliable than dirty sync data for brief outages.
|
||||
HOST1_TIER1_WRITEBACK_DELAY=60 # skip Emby writeback if outage under 1hr
|
||||
|
||||
# Tier 4 automatically syncs HOST1_DAILY_SYNC_SHARES — only list paths NOT in that array.
|
||||
FALLBACK_HOST1_WRITEBACK_TIER1=(
|
||||
"/mnt/user/Media_Server/Emby" # watch states built up during outage
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_WRITEBACK_TIER2=(
|
||||
"/mnt/user/appdata-Fallback/Important-Data" # NextCloud + Postgres — files added during outage
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_WRITEBACK_TIER3=(
|
||||
# "location-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_WRITEBACK_TIER4=(
|
||||
"/mnt/user/appdata-Fallback/Arrs_Stack" # arr databases — downloads queued during outage
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MEDIA ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Media Permissions ━━━
|
||||
# Shares that media_shares_permissions.sh applies PERMISSIONS_MODE and PERMISSIONS_OWNER to.
|
||||
# Runs first in DAILY_MAINTENANCE_SCRIPTS — arr cleanup depends on correct ownership.
|
||||
HOST1_MEDIA_PERMISSION_SHARES=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Movies-Old
|
||||
/mnt/user/Anime_Shows
|
||||
/mnt/user/Anime_Shows-Old
|
||||
/mnt/user/appcache
|
||||
/mnt/user/Books
|
||||
/mnt/user/Downloads
|
||||
/mnt/user/Games
|
||||
/mnt/user/Intros
|
||||
/mnt/user/Kids_Movies
|
||||
/mnt/user/Kids_Tv_Shows
|
||||
/mnt/user/Movie_Recordings
|
||||
/mnt/user/Movies
|
||||
/mnt/user/Music
|
||||
/mnt/user/Music_Videos
|
||||
/mnt/user/Photo
|
||||
/mnt/user/Sports
|
||||
/mnt/user/stand-up_comedy
|
||||
/mnt/user/Temp_Storage
|
||||
/mnt/user/Tv_Recordings
|
||||
/mnt/user/Tv_Shows
|
||||
/mnt/user/YouTube
|
||||
)
|
||||
|
||||
# ━━━ Media Cleaner ━━━
|
||||
# Folder lists for media_cleaner.sh — two profiles: anime and media.
|
||||
# File patterns shared across all servers — defined in master.conf.
|
||||
# Called via DAILY_MAINTENANCE_SCRIPTS. Run manually: Media/media_cleaner.sh anime|media
|
||||
HOST1_ANIME_CLEAN_FOLDERS=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Movies-Old
|
||||
/mnt/user/Anime_Shows
|
||||
/mnt/user/Anime_Shows-Old
|
||||
)
|
||||
|
||||
HOST1_MEDIA_CLEAN_FOLDERS=(
|
||||
/mnt/user/Kids_Movies
|
||||
/mnt/user/Kids_Tv_Shows
|
||||
/mnt/user/Movies
|
||||
/mnt/user/Music
|
||||
/mnt/user/Sports
|
||||
/mnt/user/stand-up_comedy
|
||||
/mnt/user/Tv_Shows
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MONITORS ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Certificate Monitor ━━━
|
||||
# Domains checked via direct openssl connection — not relying on NPM's certificate state.
|
||||
# Checks the actual certificate served, not what NPM thinks it has.
|
||||
# Thresholds (CERT_WARN_DAYS, CERT_CRIT_DAYS) defined in master.conf.
|
||||
HOST1_CERT_MONITOR_DOMAINS=(
|
||||
"Gmer4Lfe.com"
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# ━━━ SMART Health ━━━
|
||||
# Drives skipped in SMART attribute monitoring — hardware is server-specific.
|
||||
# Thresholds read from dynamix.cfg at runtime — fallbacks in master.conf.
|
||||
HOST1_SMART_IGNORE_DRIVES=(
|
||||
"sda" # boot USB — SMART not meaningful on flash drives
|
||||
)
|
||||
|
||||
# ━━━ ZFS Report ━━━
|
||||
# Pools excluded from the weekly ZFS health report — reduces noise from single-disk array pools.
|
||||
# These are individual array disks formatted as ZFS — converting to XFS over time via unBalance.
|
||||
# Pool health thresholds defined in master.conf.
|
||||
HOST1_ZFS_REPORT_IGNORE_POOLS=(
|
||||
"disk5"
|
||||
"disk6"
|
||||
"disk8"
|
||||
"disk9"
|
||||
"disk10"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── TRANSCODES ────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Ramdisk size ceiling — tmpfs only uses RAM actually needed, not the full size upfront.
|
||||
# Real-world: 9 streams peaked at ~5.5GB — 10G gives generous headroom on 128GB RAM.
|
||||
HOST1_RAMDISK_SIZE="10G"
|
||||
|
||||
# Usage thresholds — coupled to HOST1_RAMDISK_SIZE, adjust all three together if size changes.
|
||||
# Hysteresis gap (8.5 - 7 = 1.5GB) prevents flip-flop between ramdisk and SSD.
|
||||
HOST1_RAMDISK_WARN_GB=8.5 # flip to SSD when ramdisk usage reaches this
|
||||
HOST1_RAMDISK_LOW_GB=7 # flip back to ramdisk when usage drops to this
|
||||
|
||||
# SSD fallback path — where transcodes land when ramdisk exceeds HOST1_RAMDISK_WARN_GB.
|
||||
# Must be on cache pool — array disks too slow for active transcode writes.
|
||||
HOST1_TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/"
|
||||
|
||||
# Media servers sharing the ramdisk transcode space on HOST1.
|
||||
# Format: "ContainerName|URL|APIKey|Type" — Type: emby | jellyfin | plex
|
||||
# Entries with placeholder API keys are skipped automatically.
|
||||
# ⚠️ Tdarr does NOT belong here — keep Tdarr on SSD, not ramdisk.
|
||||
HOST1_TRANSCODE_SERVERS=(
|
||||
"${HOST1_EMBY_CONTAINER}|${HOST1_EMBY_URL}|${HOST1_EMBY_API_KEY}|emby"
|
||||
"${HOST1_JELLYFIN_CONTAINER}|${HOST1_JELLYFIN_URL}|${HOST1_JELLYFIN_API_KEY}|jellyfin"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── ARR STACK ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Used by arr cleanup scripts and arrs_failed_stalled_recovery.sh.
|
||||
# detect_hosts() selects HOST1 vars when running on HOST1.
|
||||
#
|
||||
# PATH MAPS — container path → host path translation.
|
||||
# Arr stores file paths using container-internal paths — scripts need host paths to scan.
|
||||
# Add one entry per root folder in arr Settings → Media Management → Root Folders.
|
||||
|
||||
# ━━━ Downloaders ━━━
|
||||
# Used by downloaders_reset.sh — runs every 30min via CRITICAL_MAINTENANCE_SCRIPTS.
|
||||
# Clears stuck states, purges old history, prepares each client for a clean cycle.
|
||||
|
||||
# slskd — clears stuck searches, dead transfers, purges expired failed imports.
|
||||
# SLSKD_FAILED_IMPORTS_DIR: where Soularr moves albums Lidarr rejected.
|
||||
HOST1_SLSKD_URL="http://localhost:8980"
|
||||
HOST1_SLSKD_API_KEY="4bF9kL2mNpQrT7vWxYz1A3dEgHjKoRsU"
|
||||
HOST1_SLSKD_FAILED_IMPORTS_DIR="/mnt/user/Temp_Storage/Slskd/completed/failed_imports"
|
||||
|
||||
# SABnzbd
|
||||
HOST1_SABNZBD_URL="http://localhost:8180"
|
||||
HOST1_SABNZBD_API_KEY="8bfefe41d83b4d50883e32859b55ca9a"
|
||||
|
||||
# qBittorrent — deleteFiles=false removes torrent from qBit but leaves files on disk.
|
||||
# Radarr/Sonarr manage actual files independently.
|
||||
HOST1_QBIT_URL="http://localhost:8080"
|
||||
HOST1_QBIT_USERNAME="root"
|
||||
HOST1_QBIT_PASSWORD="Stay0utD!ck"
|
||||
|
||||
# ━━━ Lidarr — HOST1 only ━━━
|
||||
# HOST2 does not run Lidarr — HOST1_LIDARR_RECOVERY flag handles the exit cleanly.
|
||||
HOST1_LIDARR_URL="http://localhost:8686"
|
||||
HOST1_LIDARR_API_KEY="b2977e71ef074bc0a0529d9fcce3b2dc"
|
||||
HOST1_LIDARR_MUSIC_ROOT="/mnt/user/Music-New"
|
||||
HOST1_FANART_API_KEY="Yd7147a43b692df0b364b94dc47efb81"
|
||||
HOST1_LASTFM_API_KEY="be6dc169c33ae263e690c30d18b7491d"
|
||||
|
||||
declare -A HOST1_LIDARR_PATH_MAP=(
|
||||
["/ext-music"]="/mnt/user/Music-New"
|
||||
)
|
||||
|
||||
# ━━━ Sonarr ━━━
|
||||
HOST1_SONARR_URL="http://localhost:8989"
|
||||
HOST1_SONARR_API_KEY="130decd3db5b4c25afad64864cd03f9f"
|
||||
HOST1_SONARR_TV_ROOT="/mnt/user/Tv_Shows"
|
||||
|
||||
# Note: stand-up_comedy in both Sonarr + Radarr — TV specials and movie specials, one folder
|
||||
declare -A HOST1_SONARR_PATH_MAP=(
|
||||
["/tv"]="/mnt/user/Tv_Shows"
|
||||
["/ext-standup-comedy"]="/mnt/user/stand-up_comedy"
|
||||
["/kids tv"]="/mnt/user/Kids_Tv_Shows"
|
||||
["/ext-anime-shows"]="/mnt/user/Anime_Shows-Old"
|
||||
)
|
||||
|
||||
# ━━━ Radarr ━━━
|
||||
HOST1_RADARR_URL="http://localhost:7878"
|
||||
HOST1_RADARR_API_KEY="d43a3ec6cf1549edb4af0cc63f98b2a9"
|
||||
HOST1_TMDB_API_KEY="3dac5e2e49b5540472d2eafec4f01260"
|
||||
HOST1_RADARR_MOVIES_ROOT="/mnt/user/Movies"
|
||||
|
||||
# Note: stand-up_comedy in both Radarr + Sonarr — movie specials and TV specials, one folder
|
||||
declare -A HOST1_RADARR_PATH_MAP=(
|
||||
["/movies"]="/mnt/user/Movies"
|
||||
["/kids movies"]="/mnt/user/Kids_Movies"
|
||||
["/ext-stand-up-comedy"]="/mnt/user/stand-up_comedy"
|
||||
["/anime-movies"]="/mnt/user/Anime_Movies-Old"
|
||||
)
|
||||
|
||||
# ━━━ Arr Recovery Toggles ━━━
|
||||
# false = skip that arr on this host — exits cleanly without error
|
||||
HOST1_SONARR_RECOVERY=true
|
||||
HOST1_RADARR_RECOVERY=true
|
||||
HOST1_LIDARR_RECOVERY=true # HOST1 only — exits cleanly on HOST2
|
||||
|
||||
# ==============================================================================================
|
||||
# ── SYSTEM WATCHDOG ───────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Per-host check toggles and NIC config for system_watchdog.sh.
|
||||
# Aliased by detect_hosts() — script uses unprefixed SYS_WATCHDOG_* names.
|
||||
# HOST1: TR1950X 128GB — full media server, active transcoding, ZFS cache pools.
|
||||
#
|
||||
# Three-tier response — all critical checks enabled by default on HOST1:
|
||||
# Tier 1 (bypass strikes, reboot now): docker daemon, rootfs full, kernel oops, FD, /boot
|
||||
# Tier 2 (bypass strikes with OOM): RAM critical + OOM kills in cycle
|
||||
# Tier 3 (standard strike system): everything else
|
||||
#
|
||||
# RAM tiers, OOM limits, and reboot loop settings in master.conf System Watchdog section.
|
||||
|
||||
# ━━━ Primary NIC ━━━
|
||||
# Network interface for NIC state check — verify with: ip link show | grep "^[0-9]"
|
||||
# Common values: eth0, bond0, br0, eno1
|
||||
HOST1_SYS_WATCHDOG_NIC="eth0"
|
||||
|
||||
# ━━━ Tier 1 — Critical Checks ━━━
|
||||
# These bypass the strike system — a single hit triggers immediate reboot.
|
||||
# Disabling any of these is not recommended — they protect against acute system failure.
|
||||
|
||||
# Docker daemon unresponsive → try restart, reboot if restart fails.
|
||||
# Without a working daemon docker_watchdog.sh is blind and containers cannot be managed.
|
||||
HOST1_SYS_WATCHDOG_CHECK_DOCKER_DAEMON=true
|
||||
|
||||
# rootfs at critical threshold (ROOTFS_CRITICAL_PCT=99) → reboot immediately.
|
||||
# At 99% rootfs writes fail silently — logs stop, Docker errors out, SSH may stop working.
|
||||
# Standard 95% threshold still uses strike system — only 99%+ is critical tier.
|
||||
HOST1_SYS_WATCHDOG_CHECK_ROOTFS=true
|
||||
|
||||
# Kernel BUG/Oops in dmesg delta since last cycle → reboot immediately.
|
||||
# A kernel oops means the kernel ran with a corrupted state — stability is not guaranteed.
|
||||
HOST1_SYS_WATCHDOG_CHECK_KERNEL_OOPS=true
|
||||
|
||||
# File descriptor exhaustion at FD_CRITICAL_PCT (95%) → reboot immediately.
|
||||
# At 95% FD: new connections fail, Docker can't spawn processes, SSH drops.
|
||||
HOST1_SYS_WATCHDOG_CHECK_FD=true
|
||||
|
||||
# /boot read-only detected → reboot immediately.
|
||||
# Unexpected read-only /boot means state files and config writes are silently failing.
|
||||
# Fallback state, watchdog reboot log, and lock files all go stale silently.
|
||||
HOST1_SYS_WATCHDOG_CHECK_BOOT=true
|
||||
|
||||
# ━━━ Tier 2 — Urgent OOM Check ━━━
|
||||
# Bypass strikes when RAM is critically low AND OOM kill rate confirms active crisis.
|
||||
# Both must be enabled for Tier 2 bypass to function — disable either to always use strikes.
|
||||
|
||||
# Track kernel OOM kills each cycle via /proc/vmstat oom_kill delta.
|
||||
# Also provides diagnostic context in reboot messages (which processes were killed).
|
||||
HOST1_SYS_WATCHDOG_CHECK_OOM=true
|
||||
|
||||
# Free RAM check — required for both Tier 2 bypass and RAM tier logic.
|
||||
# Tiers: MEM_WARN_GB(10) → notify | MEM_SHUTDOWN_GB(6) → stop containers | MEM_GB(4) → strikes
|
||||
HOST1_SYS_WATCHDOG_CHECK_RAM=true
|
||||
|
||||
# ━━━ Tier 3 — Standard Checks (strike system) ━━━
|
||||
# Each check must fail SYS_WATCHDOG_STRIKE_LIMIT consecutive cycles before action is taken.
|
||||
# Single spikes are ignored — sustained problems trigger reboot.
|
||||
|
||||
# /var/log filesystem usage above SYS_WATCHDOG_LOG_PCT.
|
||||
# Log spam (Docker log storms, syslog loops) fills rootfs — indicates something broken.
|
||||
HOST1_SYS_WATCHDOG_CHECK_LOG=true
|
||||
|
||||
# ZFS ARC memory pinned above SYS_WATCHDOG_ARC_PINNED_PCT after cache drop.
|
||||
# Enabled on HOST1 — ZFS cache pools actively used. Disable on hosts without ZFS.
|
||||
HOST1_SYS_WATCHDOG_CHECK_ARC=true
|
||||
|
||||
# CPU temperature above SYS_WATCHDOG_CPU_TEMP_MAX (95°C).
|
||||
# Sustained high temp causes kernel throttling or panic. Requires lm-sensors.
|
||||
HOST1_SYS_WATCHDOG_CHECK_CPU_TEMP=true
|
||||
|
||||
# Load average above SYS_WATCHDOG_LOAD_MULTIPLIER × core count.
|
||||
# DISABLED on HOST1 — Tdarr and Emby cause legitimate sustained load spikes during encoding.
|
||||
# Enable on idle servers or adjust SYS_WATCHDOG_LOAD_MULTIPLIER if load is always high.
|
||||
HOST1_SYS_WATCHDOG_CHECK_LOAD=false
|
||||
|
||||
# Zombie process count above SYS_WATCHDOG_ZOMBIE_LIMIT (50).
|
||||
# Large zombie counts indicate serious process management failure — something is stuck.
|
||||
HOST1_SYS_WATCHDOG_CHECK_ZOMBIES=true
|
||||
|
||||
# Check docker_watchdog.sh persistent skip list — required containers on skip list.
|
||||
# Cross-watchdog coordination: if docker_watchdog gave up, system_watchdog escalates.
|
||||
# ENABLED — HOST1 fully built and operational, skip list is meaningful.
|
||||
HOST1_SYS_WATCHDOG_CHECK_CONTAINERS=true
|
||||
|
||||
# /tmp filesystem usage above SYS_WATCHDOG_TMP_PCT with auto-clear attempt.
|
||||
# Script tries to clear aged /tmp files first — only strikes if clear fails.
|
||||
# Lock files, rsync temp files, and Docker ops use /tmp — 100% means lock failures.
|
||||
HOST1_SYS_WATCHDOG_CHECK_TMP=true
|
||||
|
||||
# Array disk error count delta in /proc/mdstat — accumulating errors = disk failing now.
|
||||
# Triggers on SYS_WATCHDOG_MDSTAT_ERROR_LIMIT new errors in one cycle.
|
||||
HOST1_SYS_WATCHDOG_CHECK_MDSTAT=true
|
||||
|
||||
# Primary NIC operstate — detects NIC going down (physical or driver failure).
|
||||
# Uses HOST1_SYS_WATCHDOG_NIC above. Strike system — brief flaps don't trigger reboot.
|
||||
HOST1_SYS_WATCHDOG_CHECK_NETWORK=true
|
||||
|
||||
# sshd running check — attempts restart before escalating.
|
||||
# sshd down = no remote access. Script tries rc.sshd start, notifies, strikes on failure.
|
||||
HOST1_SYS_WATCHDOG_CHECK_SSHD=true
|
||||
|
||||
# Runaway process detection — single process above SYS_WATCHDOG_RUNAWAY_CPU_PCT sustained.
|
||||
# DISABLED — Tdarr encoding and Emby transcoding legitimately peg CPU for extended periods.
|
||||
# Enable only if HOST1 has no CPU-intensive workloads.
|
||||
HOST1_SYS_WATCHDOG_CHECK_RUNAWAY=false
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RESOURCE MANAGER ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Containers to manage under pressure — see master.conf RW_CRITICAL_CONTAINERS for exclusions.
|
||||
|
||||
# docker pause at medium pressure (RAM < RW_RAM_MEDIUM_GB or load > medium threshold)
|
||||
# Suspended in-place — instant to pause/unpause, no state lost, no restart delay.
|
||||
HOST1_RW_PAUSE_CONTAINERS=(
|
||||
"Huntarr" # arr search automation — safe to suspend
|
||||
"Cleanuparr" # download cleanup — safe to suspend
|
||||
"Healarr" # arr health checks — safe to suspend
|
||||
"Soularr" # Slskd automation — background only
|
||||
"ChannelTube" # YouTube archiver — background only
|
||||
"Pinchflat" # YouTube archiver — background only
|
||||
)
|
||||
|
||||
# docker stop at hard pressure (RAM < RW_RAM_HARD_GB)
|
||||
# Full stop — these are optional/heavy services that free significant RAM when stopped.
|
||||
# resource_watchdog.sh restarts them when pressure fully clears (RAM >= RW_RAM_RECOVER_GB).
|
||||
HOST1_RW_STOP_CONTAINERS=(
|
||||
"LocalAI" # GPU/CPU heavy — largest RAM consumer when idle
|
||||
"7DaysToDie" # game server — optional
|
||||
"V-Rising" # game server — optional
|
||||
"Code-Server" # IDE — not needed during pressure events
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ──────────────────────── End Of HOST1 Variables ──────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
@@ -1,640 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================== HOST2 CONFIGURATION — unRAID-Jayred365 ===========================
|
||||
# ==============================================================================================
|
||||
# HOST2-specific variables — credentials, container names, share paths, failover lists.
|
||||
# Sourced after master.conf — values here extend shared profile arrays and add HOST2-specific
|
||||
# identity, credentials, and container configuration.
|
||||
#
|
||||
# Sparse checkout (git) ensures HOST1 never receives this file.
|
||||
# HOST1 never sees HOST2 credentials — clean separation at the file level.
|
||||
#
|
||||
# DO NOT put shared config here — thresholds, toggles, profiles belong in master.conf.
|
||||
# DO NOT put HOST1 variables here — they belong in host1.conf.
|
||||
#
|
||||
# ── STATUS ────────────────────────────────────────────────────────────────────────────────────
|
||||
# HOST2 is currently being rebuilt — most sections scaffolded, fill in when back online.
|
||||
# When ready: set FALLBACK_ENABLED=true and DAILY_RSYNC_ENABLED=true in master.conf.
|
||||
#
|
||||
# ── INDEX ─────────────────────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# ── IDENTITY & CONNECTIVITY ────────────────────────────────────────────────────────────────
|
||||
# IDENTITY hostname, SSH key
|
||||
# EMBY container name, URL, API key
|
||||
# NOTIFICATIONS Discord webhook
|
||||
# PARTNERSHIP auth containers, backup paths
|
||||
#
|
||||
# ── RSYNC ──────────────────────────────────────────────────────────────────────────────────
|
||||
# DAILY SYNC SHARES media shares HOST2 owns and pushes to HOST1
|
||||
# WEEKLY SYNC SHARES appdata shares synced weekly (Sunday 2:30am)
|
||||
# CRITICAL SYNC SHARES appdata shares synced every 30 minutes
|
||||
# BACKUP VERIFY shares for checksum verification against remote
|
||||
# HOST2 RSYNC PROFILE host2-appdata profile for HOST2-specific appdata syncs
|
||||
#
|
||||
# ── DOCKER ─────────────────────────────────────────────────────────────────────────────────
|
||||
# DOCKER DAILY RESTART containers restarted daily
|
||||
# DOCKER WEEKLY RESTART containers restarted weekly
|
||||
# DOCKER WATCHDOG memory limits, health URLs, required containers, ignore list
|
||||
# DOCKER NETWORK CONNECT networks and containers for docker_network_connect.sh
|
||||
#
|
||||
# ── FALLBACK ───────────────────────────────────────────────────────────────────────────────
|
||||
# DDNS DDNS containers managed by HOST2
|
||||
# INTERNET LOSS containers stopped when internet is lost
|
||||
# FALLBACK TIERS what HOST2 runs for HOST1 per tier
|
||||
# TIER DELAYS how long HOST2 must be down before each tier activates on HOST1
|
||||
# RSYNC WRITEBACK HOST2 appdata synced back on handback
|
||||
#
|
||||
# ── MEDIA ──────────────────────────────────────────────────────────────────────────────────
|
||||
# MEDIA PERMISSIONS share list for media_shares_permissions.sh
|
||||
# MEDIA CLEANER folder lists for media_cleaner.sh
|
||||
#
|
||||
# ── MONITORS ───────────────────────────────────────────────────────────────────────────────
|
||||
# CERTIFICATE MONITOR domains checked for SSL expiry
|
||||
# SMART HEALTH drives to skip in SMART monitoring
|
||||
# ZFS REPORT pools to exclude from ZFS health report
|
||||
#
|
||||
# ── TRANSCODES ─────────────────────────────────────────────────────────────────────────────
|
||||
# TRANSCODES ramdisk size, thresholds, SSD path, server array
|
||||
#
|
||||
# ── ARR STACK ──────────────────────────────────────────────────────────────────────────────
|
||||
# SONARR URL, API key, path map
|
||||
# RADARR URL, API key, path map
|
||||
# ARR RECOVERY per-arr recovery toggles (no Lidarr on HOST2)
|
||||
#
|
||||
# ── SYSTEM WATCHDOG ────────────────────────────────────────────────────────────────────────
|
||||
# SYSTEM WATCHDOG per-host check toggles and NIC configuration
|
||||
#
|
||||
# ── RESOURCE MANAGER ───────────────────────────────────────────────────────────────────────
|
||||
# RESOURCE MANAGER containers paused/stopped under memory pressure
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
# ==============================================================================================
|
||||
# ── IDENTITY & CONNECTIVITY ───────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Identity ━━━
|
||||
# HOST2 hostname lives in master.conf (not a credential — safe for all servers).
|
||||
# SSH key used for all server-to-server operations — rsync, failover container commands.
|
||||
# Must be in /root/.ssh/ and authorised in HOST1's /root/.ssh/authorized_keys.
|
||||
HOST2_SSH_KEY="/root/.ssh/Jayred365-rsync-key"
|
||||
HOST2_OWNER="jayred365"
|
||||
HOST2_OWNER_EMAIL="" # fill in when HOST2 is back online
|
||||
|
||||
# ━━━ Emby ━━━
|
||||
# Referenced by transcode_manager.sh, emby_session_report.sh, emby_database_repair.sh,
|
||||
# weekly_sync_maintenance.sh, and HOST2_TRANSCODE_SERVERS below.
|
||||
# API key: Emby Dashboard → API Keys → + New Key
|
||||
HOST2_EMBY_CONTAINER="Emby-Jayred365"
|
||||
HOST2_EMBY_URL="http://localhost:8096"
|
||||
HOST2_EMBY_API_KEY="your-host2-emby-api-key"
|
||||
|
||||
# ━━━ Jellyfin ━━━
|
||||
# API key: Jellyfin Dashboard → Administration → API Keys → + New Key
|
||||
HOST2_JELLYFIN_CONTAINER="Jellyfin"
|
||||
HOST2_JELLYFIN_URL="http://localhost:8095"
|
||||
HOST2_JELLYFIN_API_KEY="956d0168987f4e4680626653abb080f0"
|
||||
|
||||
# ━━━ Notifications ━━━
|
||||
# Discord webhook — leave blank to disable.
|
||||
# Per-host so HOST1 and HOST2 can post to different channels or only one server notifies.
|
||||
HOST2_DISCORD_WEBHOOK=""
|
||||
|
||||
# ━━━ Partnership ━━━
|
||||
# HOST2 is the mirror — HOST1 is always the owner unless --transfer has been run.
|
||||
# See README-Partnership.md and master.conf PARTNERSHIP section for full lifecycle docs.
|
||||
|
||||
# Auth containers reconfigured on onboard/offboard.
|
||||
# Format: "ContainerName|WebUIPort"
|
||||
# On onboard → WebUI pointed at owner's Tailscale IP (mirror clicks NPM, gets owner's auth)
|
||||
# On offboard → WebUI pointed back at localhost
|
||||
HOST2_PARTNERSHIP_AUTH_WEBUIS=(
|
||||
# fill in when HOST2 is back online
|
||||
# "NginxProxyManager|81"
|
||||
)
|
||||
|
||||
# Containers to stop on this server before the owner deploys the auth stack during onboard.
|
||||
# List whatever auth/proxy containers are currently running here.
|
||||
HOST2_PARTNERSHIP_REPLACE_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
"Authelia"
|
||||
"Authelia-Secondary"
|
||||
"Mariadb-Authelia"
|
||||
"Mariadb-Authelia-Secondary"
|
||||
"Redis-Authelia"
|
||||
"Redis-Authelia-Secondary"
|
||||
"Lldap"
|
||||
)
|
||||
|
||||
# Arr containers to stop on this server before the owner deploys the arr stack during onboard.
|
||||
HOST2_PARTNERSHIP_ARR_REPLACE_CONTAINERS=(
|
||||
# "Sonarr"
|
||||
# "Radarr"
|
||||
# "Lidarr"
|
||||
# "Prowlarr"
|
||||
# "Bazarr"
|
||||
)
|
||||
|
||||
# Paths HOST1 should collect during the grace window after offboard.
|
||||
# Notified on offboard — no auto-deletion, HOST1 must collect manually within PARTNERSHIP_GRACE_HOURS.
|
||||
HOST2_PARTNERSHIP_MIRROR_BACKUPS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# Containers parked on this server when partnership is active.
|
||||
# Stopped on onboard (owner deploys its stack instead), restarted on offboard.
|
||||
HOST2_PARTNERSHIP_OWN_CONTAINERS=(
|
||||
# "Emby"
|
||||
# "NginxProxyManager"
|
||||
)
|
||||
|
||||
# This server's desired Emby admin account on the shared Emby instance.
|
||||
# Set these — owner reads them during --onboard to create the account.
|
||||
HOST2_PARTNERSHIP_EMBY_ADMIN_USER="" # desired Emby username
|
||||
HOST2_PARTNERSHIP_EMBY_ADMIN_PASS="" # desired Emby password
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RSYNC ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Daily Sync Shares ━━━
|
||||
# Shares HOST2 pushes to all other nodes every night (1am via daily_sync_maintenance.sh).
|
||||
# Mesh model: every node pushes every media share — no ownership, no mirrors.
|
||||
# arr_sync ensures all arr libraries converge (union). rsync spreads files (additive, no --delete).
|
||||
# arr_cleanup removes true orphans based on local arr state.
|
||||
# Any node can download content to any share — it propagates to all nodes on the next cycle.
|
||||
# Nextcloud excluded — personal data, not arr-managed, synced HOST1→HOST2 only as offsite backup.
|
||||
# Uses DEFAULT_RSYNC_OPTS from master.conf — no profile needed.
|
||||
# For shares needing container stops or custom options — add a profile in master.conf.
|
||||
HOST2_DAILY_SYNC_SHARES=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Shows
|
||||
/mnt/user/Books
|
||||
/mnt/user/Intros
|
||||
/mnt/user/Kids_Movies
|
||||
/mnt/user/Kids_Tv_Shows
|
||||
/mnt/user/Movies
|
||||
/mnt/user/Music
|
||||
/mnt/user/Music_Videos
|
||||
/mnt/user/stand-up_comedy
|
||||
/mnt/user/Sports
|
||||
/mnt/user/Tv_Shows
|
||||
/mnt/user/Anime_Shows-Old
|
||||
/mnt/user/Anime_Movies-Old
|
||||
)
|
||||
|
||||
# Personal encrypted shares — synced for offsite backup, independent of media shares.
|
||||
# ZFS encrypted at dataset level — remote receives encrypted blocks, cannot read content.
|
||||
# See README-Rsync_Setup.md for ZFS encryption setup before uncommenting.
|
||||
HOST2_PERSONAL_SHARES=(
|
||||
# /mnt/user/HOST2-Personal # uncomment after creating encrypted dataset
|
||||
)
|
||||
|
||||
# ━━━ Weekly Sync Shares ━━━
|
||||
# Appdata shares synced during the weekly maintenance window (Sunday 2:30am).
|
||||
# Containers stopped both sides before sync — full clean state guaranteed.
|
||||
# Profiles drive container stops, excludes, and options — configured in master.conf RSYNC section.
|
||||
HOST2_WEEKLY_SYNC_SHARES=(
|
||||
# fill in when HOST2 is back online
|
||||
# "/mnt/user/Media_Server/Emby"
|
||||
# "/mnt/user/appdata-Fallback/Critical-Data"
|
||||
)
|
||||
|
||||
# ━━━ Intermediate Sync Shares ━━━
|
||||
# Shares synced every 4 hours by intermediate_sync_maintenance.sh.
|
||||
# Uses DEFAULT_RSYNC_OPTS (no --delete) — for sub-daily propagation of metadata or watch state.
|
||||
# Full media share sync stays in the daily window. Leave empty to skip mid-day rsync entirely.
|
||||
HOST2_INTERMEDIATE_SYNC_SHARES=(
|
||||
# fill in when HOST2 is back online
|
||||
# Example: "/mnt/user/Emby_Metadata"
|
||||
)
|
||||
|
||||
# ━━━ Critical Sync Shares ━━━
|
||||
# Appdata shares synced every 30 minutes by critical_sync_maintenance.sh.
|
||||
# Format: "/path/to/share" or "/path/to/share|profile-name"
|
||||
HOST2_CRITICAL_SYNC_SHARES=(
|
||||
# fill in when HOST2 is back online
|
||||
# "/mnt/user/appdata-Fallback/Critical-Data|critical-fallback"
|
||||
# "/mnt/user/Media_Server/Emby|emby-fallback"
|
||||
)
|
||||
|
||||
# ━━━ Backup Verify ━━━
|
||||
# Shares verified by backup_verify.sh — random file checksum comparison against remote.
|
||||
# Leave empty to use HOST2_DAILY_SYNC_SHARES automatically.
|
||||
# Sample size and minimum file size defined in master.conf.
|
||||
HOST2_BACKUP_VERIFY_SHARES=(
|
||||
# leave empty to use HOST2_DAILY_SYNC_SHARES automatically
|
||||
)
|
||||
|
||||
# ━━━ HOST2 Rsync Profile — host2-appdata ━━━
|
||||
# HOST2-specific appdata sync profile — extends the shared PROFILE_* arrays in master.conf.
|
||||
# Use for appdata unique to HOST2.
|
||||
# Shared appdata (auth stack, Emby) use dedicated profiles defined in master.conf.
|
||||
# Run manually: bash Rsync/rsync.sh /mnt/user/appdata-Fallback/HOST2-Appdata --profile=host2-appdata
|
||||
PROFILE_RSYNC_OPTS[host2-appdata]="-av --info=progress2 --bwlimit=${PROFILE_BW_LIMIT[host2-appdata]:-8000}"
|
||||
PROFILE_BW_LIMIT[host2-appdata]=8000
|
||||
PROFILE_RETRY_COUNT[host2-appdata]=3
|
||||
PROFILE_SLEEP[host2-appdata]=300
|
||||
PROFILE_CRITICAL_CONTAINER_NAMES[host2-appdata]="" # fill in when HOST2 is back online
|
||||
PROFILE_DELAYED_CONTAINERS[host2-appdata]=""
|
||||
PROFILE_CONTAINER_DELAY[host2-appdata]=5
|
||||
PROFILE_EXCLUDE_DIRS[host2-appdata]="logs *.tmp"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── DOCKER ────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Docker Daily Restart ━━━
|
||||
# Containers restarted every day via DAILY_MAINTENANCE_SCRIPTS.
|
||||
# Fill in when HOST2 is back online — add containers that degrade without daily restart.
|
||||
HOST2_DAILY_RESTART_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
# add HOST2 daily restart containers here
|
||||
)
|
||||
|
||||
# ━━━ Docker Weekly Restart ━━━
|
||||
# Less critical services restarted weekly via WEEKLY_MAINTENANCE_SCRIPTS (Sunday 2:30am).
|
||||
# Containers already stopped for weekly sync — restart adds zero extra downtime.
|
||||
HOST2_WEEKLY_RESTART_CONTAINERS=(
|
||||
# add HOST2 weekly restart containers here
|
||||
)
|
||||
|
||||
# ━━━ Docker Watchdog ━━━
|
||||
# Per-HOST2 container configuration for docker_watchdog.sh.
|
||||
# Shared thresholds and toggles live in master.conf.
|
||||
|
||||
# Memory hard limits in MB — immediate restart if exceeded.
|
||||
# Set at "container is clearly broken" not "container is busy".
|
||||
# 20GB=20480 16GB=16384 12GB=12288 10GB=10240 8GB=8192 4GB=4096 2GB=2048 1GB=1024
|
||||
declare -A HOST2_WATCHDOG_CONTAINERS=(
|
||||
["Emby"]=16384 # fill in correct limit when HOST2 is back online
|
||||
)
|
||||
|
||||
# HTTP health check URLs — checked every cycle, strike system before restart.
|
||||
# Only add containers with a meaningful web interface to check.
|
||||
declare -A HOST2_WATCHDOG_CONTAINER_URLS=(
|
||||
["Emby"]="http://localhost:8096"
|
||||
)
|
||||
|
||||
# Required containers — must always be running on HOST2.
|
||||
# Strike system before restart — repeated failures go on skip list, auto-clears on recovery.
|
||||
# Listed in dependency order — dependencies before dependents.
|
||||
HOST2_WATCHDOG_REQUIRED_CONTAINERS=(
|
||||
"NginxProxyManager"
|
||||
# add HOST2 required containers here when back online
|
||||
)
|
||||
|
||||
# Containers to skip in Tier 2 global scan — legitimately stopped or frequently restarting.
|
||||
# Watchdog leaves these alone entirely — no restart attempts, no crash loop tracking.
|
||||
HOST2_WATCHDOG_SCAN_IGNORE=(
|
||||
# add HOST2 scan ignore containers here when back online
|
||||
)
|
||||
|
||||
# Dependency ordering — skip restarting a container if its dependency is also down.
|
||||
# Prevents watchdog from restarting dependent services before their dependencies are up.
|
||||
# SPACE-SEPARATED STRINGS — converted to array at runtime.
|
||||
declare -A HOST2_WATCHDOG_DEPENDENCIES=(
|
||||
# add HOST2 dependencies here when containers are defined
|
||||
# ["Authelia"]="Mariadb-Authelia Redis-Authelia"
|
||||
)
|
||||
|
||||
# Per-container appdata growth suppress ceilings in MB.
|
||||
# ONLY needed in specific cases — growth rate detection covers all containers automatically.
|
||||
# Use when a container legitimately has large stable data and you want to suppress false-positive
|
||||
# growth alerts. Add entries here only when a container triggers warnings it shouldn't.
|
||||
declare -A HOST2_WATCHDOG_APPDATA_SIZES=(
|
||||
# add HOST2 suppress entries here only as needed
|
||||
)
|
||||
|
||||
# ━━━ Docker Network Connect ━━━
|
||||
# Containers connected to custom networks at array start by docker_network_connect.sh.
|
||||
# Networks created if they don't exist — idempotent, safe to re-run.
|
||||
HOST2_NETWORK_CONNECT_CONTAINERS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
HOST2_NETWORK_CONNECT_NETWORKS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FALLBACK ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ DDNS ━━━
|
||||
# DDNS containers HOST2 manages — started/stopped by fallback.sh per DDNS absolute rules:
|
||||
# Internet loss → stop immediately
|
||||
# Failover → HOST1 starts HOST2's DDNS as Tier 1 (before any other containers)
|
||||
# Handback → stop HOST2's DDNS on HOST1 → rsync → start containers → start local DDNS last
|
||||
HOST2_DDNS_CONTAINERS=(
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# ━━━ Internet Loss ━━━
|
||||
# Containers stopped immediately on HOST2 when internet connection is lost.
|
||||
# Prevents external-facing services from operating without connectivity.
|
||||
FALLBACK_HOST2_STOP_ON_NO_NET=(
|
||||
"Gmer4Lfe.us"
|
||||
)
|
||||
|
||||
# ━━━ Fallback Tiers — HOST2 Runs for HOST1 ━━━
|
||||
# Containers HOST2 starts when HOST1 goes down.
|
||||
# Tier 1 is always immediate — vital services cannot wait.
|
||||
# Higher tiers activate after HOST1_TIER*_DELAY minutes (set in host1.conf).
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER1=(
|
||||
"Gmer4Lfe.com"
|
||||
"Gitea" # source of truth — must be reachable even when HOST1 auth stack is down
|
||||
"Emby"
|
||||
"VaultWarden-Gmer4Lfe"
|
||||
"Dispatcharr"
|
||||
"Dispatcharr-Basic"
|
||||
"Dispatcharr-Iptv-Users"
|
||||
"ErsatzTV-Emby"
|
||||
)
|
||||
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER2=(
|
||||
"Postgres-NextCloud"
|
||||
"NextCloud"
|
||||
"PostgreSQL_Immich"
|
||||
"Immich-Gmer4Lfe"
|
||||
)
|
||||
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER3=(
|
||||
"Gitea"
|
||||
)
|
||||
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER4=(
|
||||
"Sonarr"
|
||||
"Radarr"
|
||||
"Lidarr"
|
||||
"Readarr"
|
||||
"Prowlarr"
|
||||
"Bazarr"
|
||||
"SABnzbd-Gmer4Lfe"
|
||||
"Qbittorrent-Gmer4Lfe"
|
||||
"LidaTube"
|
||||
"Pinchflat"
|
||||
"ChannelTube"
|
||||
)
|
||||
|
||||
# ━━━ Tier Delays — HOST2's Containers on HOST1 ━━━
|
||||
# How long HOST2 must be down before each tier activates on HOST1 — in minutes.
|
||||
# Tier 1 is always immediate — no delay var needed.
|
||||
HOST2_TIER2_DELAY=240 # 4 hours — productivity services
|
||||
HOST2_TIER3_DELAY=720 # 12 hours — secondary services
|
||||
HOST2_TIER4_DELAY=1440 # 24 hours — arrs + downloaders
|
||||
|
||||
# ━━━ Rsync Writeback — HOST2 Appdata Back on Handback ━━━
|
||||
# Syncs HOST2 appdata BACK to HOST2 when it comes back online after a failover.
|
||||
# Containers stopped before writeback — clean source, no competing writes.
|
||||
#
|
||||
# HOST2_TIER1_WRITEBACK_DELAY: short outages skip Tier 1 writeback — primary state
|
||||
# is more reliable than dirty sync data for brief outages.
|
||||
HOST2_TIER1_WRITEBACK_DELAY=60 # skip writeback if outage under 1hr
|
||||
|
||||
# Tier 4 automatically syncs HOST2_DAILY_SYNC_SHARES — only list paths NOT in that array.
|
||||
FALLBACK_HOST2_WRITEBACK_TIER1=(
|
||||
# "/mnt/user/appdata-Fallback/Jayred365-Emby"
|
||||
)
|
||||
|
||||
FALLBACK_HOST2_WRITEBACK_TIER2=(
|
||||
# "/mnt/user/appdata-Fallback/Jayred365-Important"
|
||||
)
|
||||
|
||||
FALLBACK_HOST2_WRITEBACK_TIER3=(
|
||||
# "location-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOST2_WRITEBACK_TIER4=(
|
||||
"/mnt/user/appdata-Fallback/Arrs_Stack" # arr databases — downloads queued during outage
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MEDIA ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Media Permissions ━━━
|
||||
# Shares that media_shares_permissions.sh applies PERMISSIONS_MODE and PERMISSIONS_OWNER to.
|
||||
# Runs first in DAILY_MAINTENANCE_SCRIPTS — arr cleanup depends on correct ownership.
|
||||
HOST2_MEDIA_PERMISSION_SHARES=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Shows
|
||||
)
|
||||
|
||||
# ━━━ Media Cleaner ━━━
|
||||
# Folder lists for media_cleaner.sh — two profiles: anime and media.
|
||||
# File patterns shared across all servers — defined in master.conf.
|
||||
# Called via DAILY_MAINTENANCE_SCRIPTS. Run manually: Media/media_cleaner.sh anime|media
|
||||
HOST2_ANIME_CLEAN_FOLDERS=(
|
||||
/mnt/user/Anime_Movies
|
||||
/mnt/user/Anime_Shows
|
||||
)
|
||||
|
||||
HOST2_MEDIA_CLEAN_FOLDERS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MONITORS ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Certificate Monitor ━━━
|
||||
# Domains checked via direct openssl connection — not relying on NPM's certificate state.
|
||||
# Checks the actual certificate served, not what NPM thinks it has.
|
||||
# Thresholds (CERT_WARN_DAYS, CERT_CRIT_DAYS) defined in master.conf.
|
||||
HOST2_CERT_MONITOR_DOMAINS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# ━━━ SMART Health ━━━
|
||||
# Drives skipped in SMART attribute monitoring — hardware is server-specific.
|
||||
# Thresholds read from dynamix.cfg at runtime — fallbacks in master.conf.
|
||||
HOST2_SMART_IGNORE_DRIVES=(
|
||||
"sda" # boot USB — SMART not meaningful on flash drives
|
||||
)
|
||||
|
||||
# ━━━ ZFS Report ━━━
|
||||
# Pools excluded from the weekly ZFS health report — reduces noise from single-disk array pools.
|
||||
# Pool health thresholds defined in master.conf.
|
||||
HOST2_ZFS_REPORT_IGNORE_POOLS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── TRANSCODES ────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Ramdisk size ceiling — tmpfs only uses RAM actually needed, not the full size upfront.
|
||||
# Adjust HOST2_RAMDISK_WARN_GB and HOST2_RAMDISK_LOW_GB together if this changes.
|
||||
HOST2_RAMDISK_SIZE="8G"
|
||||
|
||||
# Usage thresholds — coupled to HOST2_RAMDISK_SIZE, adjust all three together if size changes.
|
||||
# Hysteresis gap (6.8 - 5.5 = 1.3GB) prevents flip-flop between ramdisk and SSD.
|
||||
HOST2_RAMDISK_WARN_GB=6.8 # flip to SSD when ramdisk usage reaches this
|
||||
HOST2_RAMDISK_LOW_GB=5.5 # flip back to ramdisk when usage drops to this
|
||||
|
||||
# SSD fallback path — where transcodes land when ramdisk exceeds HOST2_RAMDISK_WARN_GB.
|
||||
# Must be on cache pool — array disks too slow for active transcode writes.
|
||||
HOST2_TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/"
|
||||
|
||||
# Media servers sharing the ramdisk transcode space on HOST2.
|
||||
# Format: "ContainerName|URL|APIKey|Type" — Type: emby | jellyfin | plex
|
||||
# Entries with placeholder API keys are skipped automatically.
|
||||
# ⚠️ Tdarr does NOT belong here — keep Tdarr on SSD, not ramdisk.
|
||||
HOST2_TRANSCODE_SERVERS=(
|
||||
"${HOST2_EMBY_CONTAINER}|${HOST2_EMBY_URL}|${HOST2_EMBY_API_KEY}|emby"
|
||||
"${HOST2_JELLYFIN_CONTAINER}|${HOST2_JELLYFIN_URL}|${HOST2_JELLYFIN_API_KEY}|jellyfin"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── ARR STACK ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Used by arr cleanup scripts and arrs_failed_stalled_recovery.sh.
|
||||
# detect_hosts() selects HOST2 vars when running on HOST2.
|
||||
# Lidarr does not run on HOST2 — HOST1_LIDARR_RECOVERY flag handles the exit cleanly.
|
||||
#
|
||||
# PATH MAPS — container path → host path translation.
|
||||
# Arr stores file paths using container-internal paths — scripts need host paths to scan.
|
||||
# Add one entry per root folder in arr Settings → Media Management → Root Folders.
|
||||
HOST2_FANART_API_KEY="Yd7147a43b692df0b364b94dc47efb81"
|
||||
HOST2_LASTFM_API_KEY="be6dc169c33ae263e690c30d18b7491d"
|
||||
|
||||
# ━━━ Sonarr ━━━
|
||||
HOST2_SONARR_URL="http://localhost:8989"
|
||||
HOST2_SONARR_API_KEY="130decd3db5b4c25afad64864cd03f9f"
|
||||
HOST2_SONARR_TV_ROOT="/mnt/user/Anime_Shows"
|
||||
|
||||
declare -A HOST2_SONARR_PATH_MAP=(
|
||||
# fill in when HOST2 is back online
|
||||
# ["/tv"]="/mnt/user/Anime_Shows"
|
||||
)
|
||||
|
||||
# ━━━ Radarr ━━━
|
||||
HOST2_RADARR_URL="http://localhost:7878"
|
||||
HOST2_RADARR_API_KEY="d43a3ec6cf1549edb4af0cc63f98b2a9"
|
||||
HOST2_RADARR_MOVIES_ROOT="/mnt/user/Anime_Movies"
|
||||
|
||||
declare -A HOST2_RADARR_PATH_MAP=(
|
||||
# fill in when HOST2 is back online
|
||||
# ["/anime-movies"]="/mnt/user/Anime_Movies"
|
||||
)
|
||||
|
||||
# ━━━ Arr Recovery Toggles ━━━
|
||||
# false = skip that arr on this host — exits cleanly without error
|
||||
HOST2_SONARR_RECOVERY=true
|
||||
HOST2_RADARR_RECOVERY=true
|
||||
# HOST2_LIDARR_RECOVERY not set — Lidarr does not run on HOST2
|
||||
|
||||
# ==============================================================================================
|
||||
# ── SYSTEM WATCHDOG ───────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Per-host check toggles and NIC config for system_watchdog.sh.
|
||||
# Aliased by detect_hosts() — script uses unprefixed SYS_WATCHDOG_* names.
|
||||
# HOST2: i5 10th gen 64GB — being rebuilt, lighter workload, no ZFS cache pools.
|
||||
#
|
||||
# Conservative defaults during rebuild — re-enable checks as HOST2 stabilises.
|
||||
# Three-tier response — all critical checks enabled regardless of rebuild state:
|
||||
# Tier 1 (bypass strikes, reboot now): docker daemon, rootfs full, kernel oops, FD, /boot
|
||||
# Tier 2 (bypass strikes with OOM): RAM critical + OOM kills in cycle
|
||||
# Tier 3 (standard strike system): selectively disabled during rebuild
|
||||
#
|
||||
# RAM tiers, OOM limits, and reboot loop settings in master.conf System Watchdog section.
|
||||
|
||||
# ━━━ Primary NIC ━━━
|
||||
# Network interface for NIC state check — verify with: ip link show | grep "^[0-9]"
|
||||
# Common values: eth0, bond0, br0, eno1
|
||||
HOST2_SYS_WATCHDOG_NIC="eth0"
|
||||
|
||||
# ━━━ Tier 1 — Critical Checks ━━━
|
||||
# All critical checks always enabled — these protect against acute failure regardless of
|
||||
# rebuild state. Disabling any is not recommended.
|
||||
|
||||
# Docker daemon unresponsive → try restart, reboot if restart fails.
|
||||
HOST2_SYS_WATCHDOG_CHECK_DOCKER_DAEMON=true
|
||||
|
||||
# rootfs at critical threshold (ROOTFS_CRITICAL_PCT=99) → reboot immediately.
|
||||
HOST2_SYS_WATCHDOG_CHECK_ROOTFS=true
|
||||
|
||||
# Kernel BUG/Oops in dmesg delta since last cycle → reboot immediately.
|
||||
HOST2_SYS_WATCHDOG_CHECK_KERNEL_OOPS=true
|
||||
|
||||
# File descriptor exhaustion at FD_CRITICAL_PCT (95%) → reboot immediately.
|
||||
HOST2_SYS_WATCHDOG_CHECK_FD=true
|
||||
|
||||
# /boot read-only detected → reboot immediately.
|
||||
HOST2_SYS_WATCHDOG_CHECK_BOOT=true
|
||||
|
||||
# ━━━ Tier 2 — Urgent OOM Check ━━━
|
||||
# Both must be enabled for Tier 2 bypass to function.
|
||||
|
||||
# Track kernel OOM kills each cycle via /proc/vmstat oom_kill delta.
|
||||
HOST2_SYS_WATCHDOG_CHECK_OOM=true
|
||||
|
||||
# Free RAM check — 64GB RAM on HOST2, tiers adjusted relative to HOST1.
|
||||
# Update master.conf SYS_WATCHDOG_MEM_* thresholds if HOST2 needs different values.
|
||||
# Currently inheriting shared master.conf values — may want lower thresholds on 64GB.
|
||||
HOST2_SYS_WATCHDOG_CHECK_RAM=true
|
||||
|
||||
# ━━━ Tier 3 — Standard Checks (strike system) ━━━
|
||||
# Several checks disabled during rebuild — enable progressively as HOST2 stabilises.
|
||||
# Each check must fail SYS_WATCHDOG_STRIKE_LIMIT consecutive cycles before action.
|
||||
|
||||
# /var/log filesystem usage above SYS_WATCHDOG_LOG_PCT.
|
||||
HOST2_SYS_WATCHDOG_CHECK_LOG=true
|
||||
|
||||
# ZFS ARC memory check.
|
||||
# DISABLED — HOST2 has no ZFS cache pools. Enable if ZFS pools are added later.
|
||||
HOST2_SYS_WATCHDOG_CHECK_ARC=false
|
||||
|
||||
# CPU temperature above SYS_WATCHDOG_CPU_TEMP_MAX (95°C).
|
||||
HOST2_SYS_WATCHDOG_CHECK_CPU_TEMP=true
|
||||
|
||||
# Load average above SYS_WATCHDOG_LOAD_MULTIPLIER × core count.
|
||||
# DISABLED — rebuild operations cause legitimate load spikes. Enable after rebuild.
|
||||
HOST2_SYS_WATCHDOG_CHECK_LOAD=false
|
||||
|
||||
# Zombie process count above SYS_WATCHDOG_ZOMBIE_LIMIT (50).
|
||||
HOST2_SYS_WATCHDOG_CHECK_ZOMBIES=true
|
||||
|
||||
# docker_watchdog.sh persistent skip list check.
|
||||
# DISABLED during rebuild — skip list may be unreliable mid-rebuild, avoid false reboots.
|
||||
# Enable once HOST2 is fully operational and docker_watchdog.sh is running stably.
|
||||
HOST2_SYS_WATCHDOG_CHECK_CONTAINERS=false
|
||||
|
||||
# /tmp filesystem usage with auto-clear attempt.
|
||||
HOST2_SYS_WATCHDOG_CHECK_TMP=true
|
||||
|
||||
# Array disk error count delta in /proc/mdstat.
|
||||
HOST2_SYS_WATCHDOG_CHECK_MDSTAT=true
|
||||
|
||||
# Primary NIC operstate — uses HOST2_SYS_WATCHDOG_NIC above.
|
||||
HOST2_SYS_WATCHDOG_CHECK_NETWORK=true
|
||||
|
||||
# sshd running check — restart attempt before escalating.
|
||||
HOST2_SYS_WATCHDOG_CHECK_SSHD=true
|
||||
|
||||
# Runaway process detection.
|
||||
# DISABLED — rebuild workloads may legitimately peg CPU. Enable after rebuild.
|
||||
HOST2_SYS_WATCHDOG_CHECK_RUNAWAY=false
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RESOURCE MANAGER ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Containers to manage under pressure — see master.conf RW_CRITICAL_CONTAINERS for exclusions.
|
||||
|
||||
# docker pause at medium pressure (RAM < RW_RAM_MEDIUM_GB or load > medium threshold)
|
||||
# Suspended in-place — instant to pause/unpause, no state lost, no restart delay.
|
||||
HOST2_RW_PAUSE_CONTAINERS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# docker stop at hard pressure (RAM < RW_RAM_HARD_GB)
|
||||
# Full stop — these are optional/heavy services that free significant RAM when stopped.
|
||||
# resource_watchdog.sh restarts them when pressure fully clears (RAM >= RW_RAM_RECOVER_GB).
|
||||
HOST2_RW_STOP_CONTAINERS=(
|
||||
# fill in when HOST2 is back online
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ──────────────────────── End Of HOST2 Variables ──────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
@@ -0,0 +1,225 @@
|
||||
# ━━━━━ DEPLOYMENT — Manual ━━━━━
|
||||
|
||||
Procedures and flag reference for the schema layer.
|
||||
For overview see README-Deployment.md. For per-script detail see the script headers.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ PROCEDURES ━━━
|
||||
|
||||
### Adding a New Configuration Variable
|
||||
|
||||
The single most common task here, and the one with the quietest failure mode if done wrong.
|
||||
|
||||
```bash
|
||||
# 1. Add it to the template — this is the versioned schema
|
||||
# Shared threshold/toggle → Deployment/master.conf.template
|
||||
# Per-host credential/path → Deployment/host.conf.template (use the HOSTN_ prefix)
|
||||
|
||||
# 2. Add it to your own live conf so you can test immediately
|
||||
# Configurations/master.conf or Configurations/host1.conf
|
||||
|
||||
# 3. Use it in the script, with a safe default
|
||||
# [[ "${NEW_THRESHOLD:-50}" -gt ... ]]
|
||||
|
||||
# 4. Commit the script change and the template change TOGETHER
|
||||
git add Deployment/master.conf.template Watchdogs/System/storage_watchdog.sh
|
||||
git commit -m "..."
|
||||
```
|
||||
|
||||
**Step 1 is the one that gets skipped.** A script merged without its template entry works on
|
||||
the machine it was written on and nowhere else — the variable is simply empty on every other
|
||||
node, and the script takes whatever branch an empty value produces. Nothing errors.
|
||||
|
||||
Placement rule, same as everywhere in the ecosystem:
|
||||
|
||||
| Kind | Goes in |
|
||||
|------|---------|
|
||||
| Threshold, toggle, profile, job list | `master.conf.template` |
|
||||
| Credential, path, container name, per-host identity | `host.conf.template` |
|
||||
|
||||
---
|
||||
|
||||
### Previewing a Schema Change Before It Lands
|
||||
|
||||
```bash
|
||||
bash Deployment/conf_upgrade.sh \
|
||||
--template Deployment/master.conf.template \
|
||||
--target Configurations/master.conf \
|
||||
--dry-run
|
||||
```
|
||||
|
||||
Prints the full ADDED / REMOVED / KEPT report and writes nothing. No root required — safe to
|
||||
run as any user, on any target, at any time.
|
||||
|
||||
Read the **REMOVED** list carefully. A key showing as REMOVED means it is in your conf but no
|
||||
longer in the template — either genuinely deprecated, or someone forgot step 1 above and the
|
||||
next pull will drop a setting you still rely on.
|
||||
|
||||
---
|
||||
|
||||
### Applying a Schema Change by Hand
|
||||
|
||||
Normally automatic via `git_pull_execute.sh`. To run it manually:
|
||||
|
||||
```bash
|
||||
bash Deployment/conf_upgrade.sh \
|
||||
--template Deployment/master.conf.template \
|
||||
--target Configurations/master.conf \
|
||||
--backup
|
||||
```
|
||||
|
||||
`--backup` writes `master.conf.bak` first. Use it — see The Recovery Gap in the README for
|
||||
why the backup matters more here than it looks.
|
||||
|
||||
For a host conf, the template must have its prefix resolved first, exactly as
|
||||
`git_pull_execute.sh` does it:
|
||||
|
||||
```bash
|
||||
TMPL=$(mktemp)
|
||||
sed "s/HOSTN_/HOST1_/g; s/REMOTE_ID/HOST2/g" Deployment/host.conf.template > "$TMPL"
|
||||
bash Deployment/conf_upgrade.sh --template "$TMPL" --target Configurations/host1.conf --backup
|
||||
rm -f "$TMPL"
|
||||
```
|
||||
|
||||
Merging the raw template without substituting `HOSTN_` would add 140 new `HOSTN_*` keys
|
||||
alongside your real `HOST1_*` ones, and mark every real key as REMOVED.
|
||||
|
||||
---
|
||||
|
||||
### Populating Credentials on a New Host
|
||||
|
||||
```bash
|
||||
bash Deployment/conf_populate.sh --dry-run # always first
|
||||
bash Deployment/conf_populate.sh
|
||||
```
|
||||
|
||||
Reads from the services actually running on this host and fills **empty** fields only.
|
||||
Existing values are never touched without `--overwrite`.
|
||||
|
||||
What it detects: arr API keys and ports from each `config.xml`, root folders from the arr
|
||||
rootFolder API, path maps from docker volume mounts, SABnzbd/slskd/qBittorrent credentials
|
||||
from their own config files, Emby/Jellyfin containers and ports, the boot device transport,
|
||||
and the default-gateway NIC.
|
||||
|
||||
Then pushes the updated conf to partners via `conf_sync.sh` so they hold the fresh keys
|
||||
immediately. `--no-push` skips that.
|
||||
|
||||
**If a field stays empty after a run,** check the output for an ambiguity warning:
|
||||
|
||||
```
|
||||
WARN: Container prefix 'authelia' is ambiguous — matches: Authelia-Secondary Authelia
|
||||
WARN: Refusing to guess. Set the container name manually in host*.conf.
|
||||
```
|
||||
|
||||
That is working as intended. Set it by hand and re-run.
|
||||
|
||||
---
|
||||
|
||||
### After Rotating an API Key
|
||||
|
||||
```bash
|
||||
bash Deployment/conf_populate.sh --overwrite --dry-run
|
||||
bash Deployment/conf_populate.sh --overwrite
|
||||
```
|
||||
|
||||
`--overwrite` replaces detected fields even when already set. This is the intended use for it
|
||||
— rotating an arr key, rebuilding a container, or repointing at a moved service.
|
||||
|
||||
Note it overwrites **every** detected field, not just the rotated one. Run the dry-run first
|
||||
and read the list.
|
||||
|
||||
---
|
||||
|
||||
### Rebuilding a Wiped Node
|
||||
|
||||
The templates are what make this possible without copying another machine's credentials.
|
||||
|
||||
```bash
|
||||
# 1. Clone the repo — templates come with it, confs do not (gitignored)
|
||||
# 2. Seed the confs from the templates
|
||||
cp Deployment/master.conf.template Configurations/master.conf
|
||||
sed "s/HOSTN_/HOST2_/g; s/REMOTE_ID/HOST1/g" Deployment/host.conf.template > Configurations/host2.conf
|
||||
|
||||
# 3. Fill in what can be detected automatically
|
||||
bash Deployment/conf_populate.sh
|
||||
|
||||
# 4. Fill in the rest by hand — anything conf_populate cannot see:
|
||||
# hostnames, DDNS containers, fallback tier lists, sync share lists,
|
||||
# partner credentials, Discord webhook
|
||||
```
|
||||
|
||||
From then on, `git_pull_execute.sh` keeps the conf in step with the template automatically.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ FLAG REFERENCE ━━━
|
||||
|
||||
### `conf_upgrade.sh`
|
||||
|
||||
| Flag | Required | What it does |
|
||||
|------|----------|-------------|
|
||||
| `--template <file>` | yes | Source of structure and new keys |
|
||||
| `--target <file>` | yes | Existing conf — source of real values, always preserved |
|
||||
| `--dry-run` | | Print the change report, write nothing. No root needed. |
|
||||
| `--backup` | | Write `<target>.bak` before installing |
|
||||
|
||||
Takes no other flags. It does not source `load_config.sh`, so `--log` and `--status` do not
|
||||
exist here — see the script header for why that is deliberate.
|
||||
|
||||
### `conf_populate.sh`
|
||||
|
||||
| Flag | What it does |
|
||||
|------|-------------|
|
||||
| `--dry-run` | Show what would be written, truncated. Changes nothing. |
|
||||
| `--overwrite` | Replace detected fields even when already set |
|
||||
| `--no-push` | Skip pushing the updated conf to partners |
|
||||
| `--log` | Verbose per-field output |
|
||||
|
||||
---
|
||||
|
||||
## ━━━ TROUBLESHOOTING ━━━
|
||||
|
||||
### A variable is empty on one node but set on another
|
||||
|
||||
The template entry is missing. Confirm:
|
||||
|
||||
```bash
|
||||
grep -n "MY_VARIABLE" Deployment/master.conf.template Deployment/host.conf.template
|
||||
```
|
||||
|
||||
No hit means the variable was added to a conf directly and never to the template, so it has
|
||||
never reached any other node. Add it to the template; the next pull propagates it.
|
||||
|
||||
### `conf_upgrade` reports a key as REMOVED that is still in use
|
||||
|
||||
Same cause, opposite direction — the key is in your conf and in a script, but not in the
|
||||
template. Add it to the template before the next pull drops it.
|
||||
|
||||
### The conf came back with different permissions
|
||||
|
||||
It should not — the target's mode and owner are copied onto the staged file before the
|
||||
rename. If they did change, check that the target existed before the run: `chmod --reference`
|
||||
silently no-ops against a missing file.
|
||||
|
||||
### A conf edit disappeared after a pull
|
||||
|
||||
Expected if the key is not in the template. `conf_upgrade` keeps values for keys that exist in
|
||||
both; a key present only in your conf is classified REMOVED and dropped. Add it to the
|
||||
template.
|
||||
|
||||
### `conf_populate` skipped a field
|
||||
|
||||
Either the service is not running, its config file was unreadable, or the container name
|
||||
prefix was ambiguous. The last case prints an explicit warning with the full match list —
|
||||
set that value by hand.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHAT THIS FOLDER DOES NOT DO ━━━
|
||||
|
||||
- **It does not create `Configurations/`.** `conf_upgrade.sh` aborts if the target does not
|
||||
exist rather than creating a partial conf. Seeding a new node is the manual step above.
|
||||
- **It does not sync confs between hosts.** That is `System_Essentials/conf_sync.sh`.
|
||||
- **It does not validate values.** It reconciles *structure*. A threshold set to nonsense
|
||||
merges through untouched — the consuming script owns validation.
|
||||
@@ -0,0 +1,211 @@
|
||||
# ━━━━━ DEPLOYMENT ━━━━━
|
||||
|
||||
The schema layer. `Configurations/*.conf` holds every value the ecosystem runs on — and is
|
||||
gitignored, because it holds credentials. This folder holds the **templates** those confs are
|
||||
built from, and the two scripts that keep the confs in step with them.
|
||||
|
||||
Two files that are versioned, and two scripts that reconcile the unversioned confs against
|
||||
them:
|
||||
|
||||
```
|
||||
Deployment/master.conf.template 349 vars ← the versioned schema
|
||||
Deployment/host.conf.template 140 vars ← per-host schema, HOSTN_-prefixed
|
||||
Deployment/conf_upgrade.sh ← template → conf, values preserved
|
||||
Deployment/conf_populate.sh ← running services → conf, empty fields only
|
||||
```
|
||||
|
||||
> **The templates are the only versioned record of what configuration exists.** Nothing else
|
||||
> in git knows that a variable is supposed to be there.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE PROBLEM THAT BUILT THIS ━━━
|
||||
|
||||
**Config Holds Secrets, So Config Cannot Be Committed**
|
||||
`master.conf` and `host*.conf` contain API keys, passwords, SSH key paths and personal
|
||||
hostnames. They are gitignored, along with their `.bak` files:
|
||||
|
||||
```
|
||||
.gitignore:4 Configurations/host*.conf
|
||||
.gitignore:5 Configurations/master.conf
|
||||
.gitignore:6 Configurations/*.bak
|
||||
```
|
||||
|
||||
That is correct and non-negotiable. But it creates a problem: if the confs are not in git,
|
||||
then **git has no idea a new setting was ever added.** A script that starts reading
|
||||
`NEW_THRESHOLD` works on the machine where it was developed and silently fails everywhere
|
||||
else, because no other node's conf has that key.
|
||||
|
||||
**A New Node Would Start With Nothing**
|
||||
Without a versioned schema there is no way to stand up a second server, or rebuild a wiped
|
||||
one, except by hand-copying a conf from a machine that already works — which means copying
|
||||
its credentials too.
|
||||
|
||||
**Hand-Editing Confs Across Nodes Does Not Scale**
|
||||
Two servers, ~490 variables between them. Adding a setting by hand means editing it on every
|
||||
node, in the right section, with the right default, without disturbing the values already
|
||||
there. Miss one and the failure surfaces days later as a script behaving differently on one
|
||||
host.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHAT THIS FOLDER DOES ━━━
|
||||
|
||||
### 🔀 Schema Merge — `conf_upgrade.sh`
|
||||
|
||||
Merges a template into an existing conf while preserving **every value the user has already
|
||||
set**. Runs automatically from `git_pull_execute.sh` after every single pull.
|
||||
|
||||
```
|
||||
Key in template only → ADDED placeholder/default, filled in once
|
||||
Key in conf only → REMOVED deprecated in this version
|
||||
Key in both → KEPT the conf's value always wins
|
||||
Comments, blank lines → from the template — structure follows the new version
|
||||
```
|
||||
|
||||
That last rule is what makes it safe to run unattended forever: the template supplies
|
||||
*structure and new keys*, never settings. Your values cannot be overwritten by a pull.
|
||||
|
||||
The live confs currently match their templates exactly — 349 and 140 variables — which is
|
||||
what a working merge looks like.
|
||||
|
||||
### 🔎 Credential Discovery — `conf_populate.sh`
|
||||
|
||||
Reads settings out of the services actually running on this host and writes them into the
|
||||
host conf: arr API keys from each `config.xml`, ports from real docker port bindings, paths
|
||||
from real volume mounts, SABnzbd/slskd/qBittorrent credentials from their own config files.
|
||||
|
||||
Only fills **empty** fields unless `--overwrite`. Manual — it is not scheduled anywhere.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ HOW A CHANGE REACHES EVERY NODE ━━━
|
||||
|
||||
```
|
||||
You add a variable
|
||||
│
|
||||
└── edit Deployment/master.conf.template ← the versioned schema
|
||||
│
|
||||
git push
|
||||
│
|
||||
└── every node: git_pull_execute.sh
|
||||
│
|
||||
└── conf_upgrade.sh --template ... --target ... --backup
|
||||
│
|
||||
ADDED → new key appears with the template default
|
||||
KEPT → every existing value untouched
|
||||
REMOVED → deprecated keys dropped
|
||||
```
|
||||
|
||||
**This is the rule that follows from it, and it is not optional:**
|
||||
|
||||
> Any conf variable change — add, remove, or rename — must update
|
||||
> `Deployment/master.conf.template` and `Deployment/host.conf.template` **in the same pass**
|
||||
> as the script change that uses it.
|
||||
|
||||
A script merged without its template entry works only on the machine it was written on.
|
||||
Nothing errors; the variable is simply empty everywhere else, and the script takes whatever
|
||||
branch an empty value leads to.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE `HOSTN_` PLACEHOLDER ━━━
|
||||
|
||||
`host.conf.template` is written with a generic prefix — 149 occurrences of `HOSTN_`:
|
||||
|
||||
```bash
|
||||
HOSTN_SONARR_URL=""
|
||||
HOSTN_SONARR_API_KEY=""
|
||||
```
|
||||
|
||||
`git_pull_execute.sh` substitutes the real identity before merging, so keys match the target:
|
||||
|
||||
```bash
|
||||
sed "s/HOSTN_/${MY_ID}_/g; s/REMOTE_ID/${REMOTE_ID}/g" host.conf.template > "$TMPL_RESOLVED"
|
||||
```
|
||||
|
||||
One template therefore serves every host. HOST1 merges it as `HOST1_*`, HOST2 as `HOST2_*`,
|
||||
and a third node would work with no template change at all.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SCRIPTS IN THIS FOLDER ━━━
|
||||
|
||||
| Script | Role | When It Runs |
|
||||
|--------|------|-------------|
|
||||
| `conf_upgrade.sh` | Merge template into conf — structure forward, values preserved | Automatically, after every `git pull` |
|
||||
| `conf_populate.sh` | Detect settings from running services into the host conf | Manually — onboarding, or after a key rotation |
|
||||
| `migrate_data_layout.sh` | Move everything persisted into the rooted `data/` tree | Once per host, manually. Idempotent. |
|
||||
|
||||
### 📦 Data layout migration — `migrate_data_layout.sh`
|
||||
|
||||
`conf_upgrade.sh` adds keys the template has and the installation does not; it never rewrites a
|
||||
value you already have. That is exactly what you want from it, and exactly why it cannot perform
|
||||
a layout migration — the paths being moved are *existing* keys, so their values would keep
|
||||
pointing at the old layout forever while the new directory variables sat beside them unused.
|
||||
|
||||
So this rewrites those values and moves the files to match. Both halves or neither.
|
||||
|
||||
```bash
|
||||
Deployment/migrate_data_layout.sh --dry-run # always first
|
||||
Deployment/migrate_data_layout.sh
|
||||
```
|
||||
|
||||
It refuses to run while a job from *this* installation is active — scoped to the installation's
|
||||
own path, because `pgrep` is system-wide and a box running both a production checkout and a
|
||||
development clone will otherwise always look busy. `--force` overrides.
|
||||
|
||||
Each host runs it itself: `data/` is gitignored, so a restructure travels as code and conf while
|
||||
the files stay where they are. See `data/README.md` for the resulting layout.
|
||||
|
||||
| Template | Role |
|
||||
|----------|------|
|
||||
| `master.conf.template` | Shared schema — thresholds, toggles, profiles, orchestrator job lists |
|
||||
| `host.conf.template` | Per-host schema — credentials, paths, container names. `HOSTN_`-prefixed |
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SAFEGUARDS WORTH KNOWING ━━━
|
||||
|
||||
**The install is atomic.** `conf_upgrade.sh` stages the merged conf beside the target and
|
||||
installs it with a rename, never a copy. A `cp` truncates the live conf and writes into it —
|
||||
and every watchdog sources `load_config.sh` on every run, so anything reading during that
|
||||
window would get a partial conf with empty path variables. The temp file is staged in the
|
||||
target's own directory deliberately: `/tmp` is rootfs while the confs are on flash, and a
|
||||
cross-device `mv` degrades to copy-then-unlink, which is the exact torn write being avoided.
|
||||
|
||||
**Dry run needs no privilege, writing does.** `--dry-run` prints the full ADDED / REMOVED /
|
||||
KEPT report and is useful to anyone. Installing over a conf under `/boot` requires root.
|
||||
|
||||
**`conf_upgrade.sh` sources nothing — deliberately.** No `load_config.sh`, no `common.sh`.
|
||||
It is the tool that repairs the conf `load_config.sh` depends on, so it has to work when that
|
||||
conf is broken, partial, or missing keys. That is also why it uses plain `echo` rather than
|
||||
`log()`, and why it has no `acquire_lock` — concurrency is handled by the atomic rename
|
||||
instead, and since the merge is idempotent, last-writer-wins is identical to running once.
|
||||
|
||||
**`conf_populate.sh` refuses to guess a container.** An ambiguous name prefix skips the field
|
||||
rather than picking the first match. Writing the wrong container name is worse than writing
|
||||
nothing: an empty field is visibly incomplete and gets fixed, a wrong one silently points the
|
||||
whole stack at the wrong instance. This host has a live example — `authelia` prefix-matches
|
||||
both `Authelia` (9091) and `Authelia-Secondary` (9092).
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE RECOVERY GAP ━━━
|
||||
|
||||
Confs are gitignored, **and so are their `.bak` files**. There is no versioned history to
|
||||
revert to, and the single `.bak` slot is overwritten by whoever writes next:
|
||||
|
||||
| file | modified | its `.bak` |
|
||||
|---|---|---|
|
||||
| `master.conf` | Jul 28 18:52 | Jul 28 18:52 |
|
||||
| `host1.conf` | Aug 1 21:00 | **Jul 3 17:46** |
|
||||
|
||||
A bad write to `master.conf` currently falls back to a file that may predate weeks of edits.
|
||||
Worth knowing before hand-editing a conf, and the reason `--backup` exists on `conf_upgrade.sh`
|
||||
at all.
|
||||
|
||||
Moving `Configurations/` into a private repo would make `git diff` and `git revert` the
|
||||
recovery mechanism and give the history for free. That overlaps the existing GitHub-mirror
|
||||
TODO, which is blocked on the same question — see `Notes_AI-Design.md`, where it also blocks
|
||||
AI-assisted conf writes.
|
||||
Executable
+559
@@ -0,0 +1,559 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================= Conf Auto-Populate =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Reads credentials and settings from locally running services and writes
|
||||
# them into the local host conf. Safe to run multiple times — only populates
|
||||
# EMPTY fields, never overwrites existing values unless --overwrite is passed.
|
||||
#
|
||||
# After populating, pushes the updated conf to all partners via conf_sync.sh
|
||||
# so they have the fresh keys in their /tmp/varaverk/conf/ cache immediately.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# AUTO-DETECTED FIELDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# HOST CONF (Configurations/<hostid>.conf)
|
||||
# ──────────────────────────────────────────────────────────────────────────
|
||||
# HOSTN_OWNER from hostname (strip unRAID- prefix, lowercase)
|
||||
# HOSTN_SSH_KEY from hostname convention (/root/.ssh/<owner>_rsync_automation)
|
||||
# HOSTN_STORAGE_MODE_INTERNAL from boot device transport (NVMe/SSD=true, USB=false)
|
||||
# HOSTN_RADARR_API_KEY from Radarr config.xml (found via docker volume mount)
|
||||
# HOSTN_RADARR_URL from Radarr config.xml port
|
||||
# HOSTN_RADARR_MOVIES_ROOT from Radarr rootFolder API
|
||||
# HOSTN_RADARR_PATH_MAP from Radarr docker volume mounts vs root folder path
|
||||
# HOSTN_SONARR_API_KEY from Sonarr config.xml
|
||||
# HOSTN_SONARR_URL from Sonarr config.xml port
|
||||
# HOSTN_SONARR_TV_ROOT from Sonarr rootFolder API
|
||||
# HOSTN_SONARR_PATH_MAP from Sonarr docker volume mounts vs root folder path
|
||||
# HOSTN_LIDARR_API_KEY from Lidarr config.xml
|
||||
# HOSTN_LIDARR_URL from Lidarr config.xml port
|
||||
# HOSTN_LIDARR_MUSIC_ROOT from Lidarr rootFolder API
|
||||
# HOSTN_LIDARR_PATH_MAP from Lidarr docker volume mounts vs root folder path
|
||||
# HOSTN_SABNZBD_API_KEY from sabnzbd.ini
|
||||
# HOSTN_SABNZBD_URL from sabnzbd.ini port
|
||||
# HOSTN_SLSKD_API_KEY from slskd config.yml
|
||||
# HOSTN_SLSKD_URL from slskd config.yml port
|
||||
# HOSTN_QBIT_URL from qBittorrent.conf WebUI port
|
||||
# HOSTN_QBIT_USERNAME from qBittorrent.conf WebUI username
|
||||
# HOSTN_QBIT_PASSWORD from qBittorrent.conf WebUI password (plaintext only)
|
||||
# HOSTN_EMBY_CONTAINER fuzzy match from docker ps
|
||||
# HOSTN_EMBY_URL from docker port binding
|
||||
# HOSTN_JELLYFIN_CONTAINER fuzzy match from docker ps
|
||||
# HOSTN_JELLYFIN_URL from docker port binding
|
||||
# HOSTN_TRANSCODE_SSD from Emby/Jellyfin container /transcode volume mount
|
||||
# HOSTN_AUTHELIA_CONTAINER fuzzy match from docker ps
|
||||
# HOSTN_AUTHELIA_CONFIG from Authelia container /config volume mount
|
||||
# HOSTN_SYS_WATCHDOG_NIC from ip route default gateway interface
|
||||
#
|
||||
# MASTER CONF (Configurations/master.conf)
|
||||
# ──────────────────────────────────────────────────────────────────────────
|
||||
# HOST1 / HOST2 local hostname written to MY_ID slot
|
||||
# GITEA_CONTAINER fuzzy match from docker ps
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Resolve this host's conf from MY_ID
|
||||
# 2. Per service, locate its container and read its own config file:
|
||||
# arrs → config.xml via the container's /config volume mount
|
||||
# SABnzbd → sabnzbd.ini
|
||||
# slskd → config.yml
|
||||
# qBit → qBittorrent.conf (plaintext WebUI credentials only)
|
||||
# Emby/JF → docker port bindings and /transcode mount
|
||||
# 3. Write each value only if the conf field is EMPTY, unless --overwrite
|
||||
# 4. Push the updated conf to partners via conf_sync.sh, unless --no-push
|
||||
#
|
||||
# Container names are resolved by _resolve_container(): an exact name match wins, otherwise
|
||||
# a prefix match must be unambiguous or the field is skipped.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Never Overwrite a Human's Value
|
||||
# Only empty fields are populated. A value already in the conf was either set deliberately
|
||||
# or populated from a service that has since changed — either way, the file wins over
|
||||
# detection. --overwrite exists for deliberate re-sync after a key rotation.
|
||||
#
|
||||
# Read From the Service, Not From Assumption
|
||||
# Every value comes out of the service's own config file or docker metadata — ports from
|
||||
# actual port bindings, paths from actual volume mounts. Nothing is derived from naming
|
||||
# convention where the real value is readable.
|
||||
#
|
||||
# Refuse to Guess a Container
|
||||
# An ambiguous prefix skips the field rather than picking one. Writing the wrong container
|
||||
# name is worse than writing nothing: an empty field is visibly incomplete and gets fixed,
|
||||
# while a wrong one silently points the whole stack at the wrong instance. This host has a
|
||||
# live example — "authelia" prefix-matches both Authelia and Authelia-Secondary.
|
||||
#
|
||||
# Push Immediately After Populating
|
||||
# Fresh credentials go to partners right away rather than waiting for the next scheduled
|
||||
# conf sync, so a partner is never authenticating with a key this host has already rotated.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Reads service config files owned by container users and writes the host conf.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() resolves MY_ID, which selects which host conf is written. Populating the
|
||||
# wrong host's conf would write this machine's credentials into a partner's file.
|
||||
#
|
||||
# Conf Existence Guard
|
||||
# Aborts if the resolved host conf does not exist, rather than creating a partial one.
|
||||
#
|
||||
# Empty-Field-Only Writes
|
||||
# Existing values are preserved unless --overwrite is passed explicitly.
|
||||
#
|
||||
# Container Ambiguity Guard
|
||||
# _resolve_container() refuses a prefix matching more than one container, warning with the
|
||||
# full match list. Exact name matches short-circuit and are never treated as ambiguous.
|
||||
#
|
||||
# Missing Service Tolerance
|
||||
# A service that is not installed on this host is skipped with a log line. Absence is a
|
||||
# valid configuration, not a failure.
|
||||
#
|
||||
# Map Block Validation
|
||||
# Associative-array entries are only inserted when the target map actually exists in the
|
||||
# conf; a missing map warns and skips instead of appending an orphaned entry.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run reports every value it would write, truncated, and writes nothing.
|
||||
#
|
||||
# Credential Truncation in Output
|
||||
# Detected secrets are printed truncated (first 8 chars) so a populate run can be pasted
|
||||
# into a log or issue without leaking full API keys.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Reads and writes: Configurations/<hostid>.conf and Configurations/master.conf
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# DOCKER_APPDATA_BASE
|
||||
# Fallback appdata root when a container exposes no /config mount to read from.
|
||||
#
|
||||
# Everything else this script touches is a field it populates rather than one it consumes —
|
||||
# see AUTO-DETECTED FIELDS above for the full list.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# conf_populate.sh Populate empty fields only
|
||||
# conf_populate.sh --overwrite Overwrite all detected fields (re-sync after arr key rotation)
|
||||
# conf_populate.sh --dry-run Show what would be written, no changes
|
||||
# conf_populate.sh --log Verbose output
|
||||
# conf_populate.sh --no-push Skip pushing to partners after update
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
# Deployment/ sits one level under the repo root, not three. The old ../../../ resolved to
|
||||
# /boot/config on a flash install and /mnt/user on an appdata one — outside the repo either
|
||||
# way, so this sourced nothing and every helper below was "command not found". Onboard Step 11
|
||||
# has been failing on that since it was written.
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
SCRIPTS_ROOT="$SCRIPTS_DIR"
|
||||
|
||||
OVERWRITE=false
|
||||
NO_PUSH=false
|
||||
FILTERED_ARGS=()
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--overwrite) OVERWRITE=true ;;
|
||||
--no-push) NO_PUSH=true ;;
|
||||
*) FILTERED_ARGS+=("$arg") ;;
|
||||
esac
|
||||
done
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
detect_hosts
|
||||
|
||||
[[ "$EUID" -ne 0 ]] && { error "Must be run as root"; exit 1; }
|
||||
|
||||
CONF_FILE="$SCRIPTS_ROOT/Configurations/${MY_ID,,}.conf"
|
||||
[[ ! -f "$CONF_FILE" ]] && { error "Conf file not found: $CONF_FILE"; exit 1; }
|
||||
|
||||
MASTER_CONF="$SCRIPTS_ROOT/Configurations/master.conf"
|
||||
|
||||
log "$ICON_GEAR Config: conf=${CONF_FILE} overwrite=${OVERWRITE:-false} no-push=${NO_PUSH:-false}"
|
||||
|
||||
UPDATED=0
|
||||
SKIPPED=0
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR Conf Auto-Populate — $MY_ID ($LOCAL_SERVER_NAME) ━━━"
|
||||
echo ""
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
||||
[[ "$OVERWRITE" == true ]] && warn "OVERWRITE mode — existing values will be replaced"
|
||||
|
||||
# ── Helper: write a var into a conf file if empty (or --overwrite) ────────────
|
||||
# Optional 4th arg: target file (defaults to $CONF_FILE)
|
||||
_set_conf_var() {
|
||||
local var_name="$1" value="$2" label="$3" target="${4:-$CONF_FILE}"
|
||||
[[ -z "$value" ]] && return
|
||||
|
||||
local current
|
||||
current=$(grep -oP "(?<=^\s*${var_name}=\")[^\"]*" "$target" 2>/dev/null | head -1)
|
||||
|
||||
if [[ -n "$current" ]] && [[ "$OVERWRITE" == false ]]; then
|
||||
log "$label: already set (${current:0:8}…) — skipping"
|
||||
(( SKIPPED++ ))
|
||||
return
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would set $var_name = ${value:0:8}…"
|
||||
return
|
||||
fi
|
||||
|
||||
if grep -q "^\s*${var_name}=" "$target"; then
|
||||
sed -i "s|^\(\s*${var_name}\s*=\s*\)\"[^\"]*\"|\1\"${value}\"|" "$target"
|
||||
else
|
||||
printf '\n %s="%s"\n' "$var_name" "$value" >> "$target"
|
||||
fi
|
||||
info "$label: set ✅"
|
||||
(( UPDATED++ ))
|
||||
}
|
||||
|
||||
# ── Helper: add or update an entry in a declare -A map block ──────────────────
|
||||
# Inserts [key]="value" before the closing ) of the named map in $CONF_FILE.
|
||||
# Skips if an identical uncommented entry already exists (unless --overwrite).
|
||||
_set_conf_map_entry() {
|
||||
local map_name="$1" key="$2" value="$3" label="$4"
|
||||
[[ -z "$key" || -z "$value" ]] && return
|
||||
|
||||
if grep -qP "^\s*\[\"${key//\//\\/}\"\]=" "$CONF_FILE" 2>/dev/null; then
|
||||
if [[ "$OVERWRITE" == false ]]; then
|
||||
log "$label: already set — skipping"
|
||||
(( SKIPPED++ ))
|
||||
return
|
||||
fi
|
||||
[[ "$DRY_RUN" == true ]] && { warn "DRY RUN — would update ${map_name}[\"${key}\"]"; return; }
|
||||
sed -i "s|^\(\s*\)\[\"${key}\"\]=\"[^\"]*\"|\1[\"${key}\"]=\"${value}\"|" "$CONF_FILE"
|
||||
info "$label: set ✅"
|
||||
(( UPDATED++ ))
|
||||
return
|
||||
fi
|
||||
|
||||
if ! grep -q "declare -A ${map_name}=" "$CONF_FILE" 2>/dev/null; then
|
||||
warn "$label: ${map_name} not found in conf — skipping"
|
||||
return
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && { warn "DRY RUN — would set ${map_name}[\"${key}\"] = \"${value}\""; return; }
|
||||
|
||||
awk -v mapname="${map_name}" -v key="$key" -v val="$value" '
|
||||
BEGIN { in_map=0; done=0 }
|
||||
!done && index($0, "declare -A " mapname) { in_map=1 }
|
||||
in_map && !done && /^\s*\)/ {
|
||||
printf " [\"%s\"]=\"%s\"\n", key, val
|
||||
done=1; in_map=0
|
||||
}
|
||||
{ print }
|
||||
' "$CONF_FILE" > "${CONF_FILE}.tmp" && mv "${CONF_FILE}.tmp" "$CONF_FILE" || {
|
||||
rm -f "${CONF_FILE}.tmp"
|
||||
warn "$label: failed to update ${map_name}"
|
||||
return 1
|
||||
}
|
||||
info "$label: set ✅"
|
||||
(( UPDATED++ ))
|
||||
}
|
||||
|
||||
# ── Helper: resolve a container name from a prefix, refusing ambiguity ────────
|
||||
# grep -m1 silently returns whichever name docker happens to list first. On this very host
|
||||
# "^authelia" matches both Authelia (9091, primary) and Authelia-Secondary (9092), and -m1
|
||||
# picks the secondary — writing the wrong container into the conf that every other script
|
||||
# then trusts. An exact match wins outright; otherwise a prefix match must be unambiguous.
|
||||
# Same rule detect_hosts() applies to host identity: exactly one candidate, or none.
|
||||
_resolve_container() {
|
||||
local pattern="$1"
|
||||
local all exact matches count
|
||||
|
||||
all=$(docker ps -a --format '{{.Names}}' 2>/dev/null)
|
||||
[[ -z "$all" ]] && return 1
|
||||
|
||||
# Exact name match short-circuits — "Authelia" is not ambiguous with "Authelia-Secondary"
|
||||
exact=$(printf '%s\n' "$all" | grep -ixm1 -- "$pattern")
|
||||
[[ -n "$exact" ]] && { echo "$exact"; return 0; }
|
||||
|
||||
matches=$(printf '%s\n' "$all" | grep -i -- "^${pattern}")
|
||||
count=$(printf '%s\n' "$matches" | grep -c .)
|
||||
|
||||
if [[ "$count" -gt 1 ]]; then
|
||||
warn "Container prefix '${pattern}' is ambiguous — matches: $(printf '%s' "$matches" | tr '\n' ' ')"
|
||||
warn " Refusing to guess. Set the container name manually in host*.conf."
|
||||
return 1
|
||||
fi
|
||||
|
||||
[[ "$count" -eq 1 ]] && { printf '%s\n' "$matches"; return 0; }
|
||||
return 1
|
||||
}
|
||||
|
||||
# ── Helper: find arr config dir via docker volume mount ───────────────────────
|
||||
_arr_config_dir() {
|
||||
local pattern="$1"
|
||||
local container_name
|
||||
container_name=$(_resolve_container "$pattern") || return 1
|
||||
[[ -z "$container_name" ]] && return 1
|
||||
|
||||
local config_path
|
||||
config_path=$(docker inspect "$container_name" 2>/dev/null | \
|
||||
jq -r '.[0].Mounts[]? | select(.Destination == "/config") | .Source' 2>/dev/null | head -1)
|
||||
[[ -z "$config_path" ]] && config_path="${DOCKER_APPDATA_BASE:-/mnt/user/appdata}/${container_name}"
|
||||
|
||||
[[ -d "$config_path" ]] && echo "$config_path" || return 1
|
||||
}
|
||||
|
||||
# ── Helper: read XML tag value ────────────────────────────────────────────────
|
||||
_xml_val() {
|
||||
local file="$1" tag="$2"
|
||||
grep -oP "(?<=<${tag}>)[^<]+" "$file" 2>/dev/null | head -1
|
||||
}
|
||||
|
||||
# ── Helper: get host-side port for a container's internal port ────────────────
|
||||
_docker_host_port() {
|
||||
local container="$1" container_port="$2"
|
||||
docker inspect "$container" 2>/dev/null | \
|
||||
jq -r --arg p "${container_port}/tcp" \
|
||||
'.[0].NetworkSettings.Ports[$p]?[0].HostPort // empty' 2>/dev/null | head -1
|
||||
}
|
||||
|
||||
# ── Helper: get host path for a container destination mount ──────────────────
|
||||
_docker_volume_host() {
|
||||
local container="$1" dest="$2"
|
||||
docker inspect "$container" 2>/dev/null | \
|
||||
jq -r --arg d "$dest" \
|
||||
'.[0].Mounts[]? | select(.Destination == $d) | .Source' 2>/dev/null | head -1
|
||||
}
|
||||
|
||||
# ── Helper: find host↔container path mapping for a given container path ───────
|
||||
# Finds the most specific mount whose destination is a prefix of container_path
|
||||
# and where source != destination (i.e. an actual remapping exists).
|
||||
# Outputs "source|destination" or nothing if no remapping found.
|
||||
_docker_path_map() {
|
||||
local container="$1" container_path="$2"
|
||||
docker inspect "$container" 2>/dev/null | \
|
||||
jq -r '.[0].Mounts[]? | select(.Destination != "/config") | "\(.Source)|\(.Destination)"' \
|
||||
2>/dev/null | \
|
||||
awk -F'|' -v target="$container_path" '
|
||||
$1 != $2 && length($2) > 0 && index(target, $2) == 1 { print length($2), $0 }
|
||||
' | sort -rn | head -1 | cut -d' ' -f2-
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Owner short name + SSH key ────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
owner=$(echo "$LOCAL_SERVER_NAME" | sed 's/^[Uu][Nn][Rr][Aa][Ii][Dd]-//i' | tr '[:upper:]' '[:lower:]')
|
||||
_set_conf_var "${MY_ID}_OWNER" "$owner" "Owner short name"
|
||||
_set_conf_var "${MY_ID}_SSH_KEY" "/root/.ssh/${owner}_rsync_automation" "SSH key path"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Arr API keys + URLs + root paths + path maps ──────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
for arr in radarr sonarr lidarr; do
|
||||
arr_upper="${arr^^}"
|
||||
arr_container=$(_resolve_container "$arr")
|
||||
config_dir=$(_arr_config_dir "$arr") || {
|
||||
log "${arr_upper}: no running container found — skipping"
|
||||
continue
|
||||
}
|
||||
config_xml="${config_dir}/config.xml"
|
||||
|
||||
if [[ ! -f "$config_xml" ]]; then
|
||||
log "${arr_upper}: config.xml not found at $config_xml — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
key=$(_xml_val "$config_xml" "ApiKey")
|
||||
port=$(_xml_val "$config_xml" "Port")
|
||||
port="${port:-$(case $arr in radarr) echo 7878;; sonarr) echo 8989;; lidarr) echo 8686;; esac)}"
|
||||
url_base="http://localhost:${port}"
|
||||
|
||||
_set_conf_var "${MY_ID}_${arr_upper}_API_KEY" "$key" "${arr_upper} API key"
|
||||
_set_conf_var "${MY_ID}_${arr_upper}_URL" "$url_base" "${arr_upper} URL"
|
||||
|
||||
if [[ -n "$key" ]]; then
|
||||
case "$arr" in lidarr) api_ver="v1" ;; *) api_ver="v3" ;; esac
|
||||
root_json=$(curl -sf --max-time 5 \
|
||||
-H "X-Api-Key: $key" "${url_base}/api/${api_ver}/rootfolder" 2>/dev/null)
|
||||
root_path=$(echo "$root_json" | jq -r '.[0].path // empty' 2>/dev/null)
|
||||
|
||||
case "$arr" in
|
||||
radarr) _set_conf_var "${MY_ID}_RADARR_MOVIES_ROOT" "$root_path" "Radarr movies root" ;;
|
||||
sonarr) _set_conf_var "${MY_ID}_SONARR_TV_ROOT" "$root_path" "Sonarr TV root" ;;
|
||||
lidarr) _set_conf_var "${MY_ID}_LIDARR_MUSIC_ROOT" "$root_path" "Lidarr music root" ;;
|
||||
esac
|
||||
|
||||
if [[ -n "$root_path" && -n "$arr_container" ]]; then
|
||||
map_entry=$(_docker_path_map "$arr_container" "$root_path")
|
||||
if [[ -n "$map_entry" ]]; then
|
||||
map_src="${map_entry%%|*}"
|
||||
map_dest="${map_entry##*|}"
|
||||
case "$arr" in
|
||||
radarr) _set_conf_map_entry "${MY_ID}_RADARR_PATH_MAP" "$map_dest" "$map_src" "Radarr path map" ;;
|
||||
sonarr) _set_conf_map_entry "${MY_ID}_SONARR_PATH_MAP" "$map_dest" "$map_src" "Sonarr path map" ;;
|
||||
lidarr) _set_conf_map_entry "${MY_ID}_LIDARR_PATH_MAP" "$map_dest" "$map_src" "Lidarr path map" ;;
|
||||
esac
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ── SABnzbd API key + URL ─────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
sab_dir=$(_arr_config_dir "sabnzbd") && {
|
||||
sab_ini=$(find "$sab_dir" -maxdepth 2 -name "sabnzbd.ini" 2>/dev/null | head -1)
|
||||
if [[ -f "$sab_ini" ]]; then
|
||||
sab_key=$(grep -oP '(?<=^api_key\s*=\s*)\S+' "$sab_ini" 2>/dev/null | head -1)
|
||||
sab_port=$(grep -oP '(?<=^port\s*=\s*)\d+' "$sab_ini" 2>/dev/null | head -1)
|
||||
_set_conf_var "${MY_ID}_SABNZBD_API_KEY" "$sab_key" "SABnzbd API key"
|
||||
_set_conf_var "${MY_ID}_SABNZBD_URL" "http://localhost:${sab_port:-8080}" "SABnzbd URL"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── slskd API key + URL ───────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
slskd_dir=$(_arr_config_dir "slskd") && {
|
||||
slskd_yml=$(find "$slskd_dir" -maxdepth 2 \( -name "*.yml" -o -name "*.yaml" \) 2>/dev/null | head -1)
|
||||
if [[ -f "$slskd_yml" ]]; then
|
||||
slskd_key=$(grep -oP '(?<=api_key:\s)[\w-]+' "$slskd_yml" 2>/dev/null | head -1)
|
||||
[[ -z "$slskd_key" ]] && \
|
||||
slskd_key=$(grep -oP '(?<=apikey:\s)[\w-]+' "$slskd_yml" 2>/dev/null | head -1)
|
||||
slskd_port=$(grep -oP '(?<=port:\s)\d+' "$slskd_yml" 2>/dev/null | head -1)
|
||||
_set_conf_var "${MY_ID}_SLSKD_API_KEY" "$slskd_key" "slskd API key"
|
||||
_set_conf_var "${MY_ID}_SLSKD_URL" "http://localhost:${slskd_port:-5030}" "slskd URL"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── qBittorrent URL + credentials ────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
qbit_dir=$(_arr_config_dir "qbittorrent") && {
|
||||
qbit_conf=$(find "$qbit_dir" -maxdepth 3 -name "qBittorrent.conf" 2>/dev/null | head -1)
|
||||
if [[ -f "$qbit_conf" ]]; then
|
||||
qbit_port=$(grep -oP '(?<=WebUI\\Port=)\d+' "$qbit_conf" 2>/dev/null | head -1)
|
||||
qbit_user=$(grep -oP '(?<=WebUI\\Username=)\S+' "$qbit_conf" 2>/dev/null | head -1)
|
||||
# Only capture plaintext password — PBKDF2 hashes are not usable
|
||||
qbit_pass=$(grep -oP '(?<=WebUI\\Password=)[^\r\n]+' "$qbit_conf" 2>/dev/null | \
|
||||
grep -v '@ByteArray' | head -1)
|
||||
_set_conf_var "${MY_ID}_QBIT_URL" "http://localhost:${qbit_port:-8080}" "qBittorrent URL"
|
||||
_set_conf_var "${MY_ID}_QBIT_USERNAME" "$qbit_user" "qBittorrent username"
|
||||
_set_conf_var "${MY_ID}_QBIT_PASSWORD" "$qbit_pass" "qBittorrent password"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Media server container names + URLs + transcode path ─────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
transcode_dir=""
|
||||
|
||||
for pattern in "emby" "jellyfin"; do
|
||||
container=$(_resolve_container "$pattern") || continue
|
||||
[[ -z "$container" ]] && continue
|
||||
|
||||
case "$pattern" in
|
||||
emby)
|
||||
host_port=$(_docker_host_port "$container" "8096")
|
||||
_set_conf_var "${MY_ID}_EMBY_CONTAINER" "$container" "Emby container name"
|
||||
_set_conf_var "${MY_ID}_EMBY_URL" "http://localhost:${host_port:-8096}" "Emby URL"
|
||||
;;
|
||||
jellyfin)
|
||||
host_port=$(_docker_host_port "$container" "8096")
|
||||
_set_conf_var "${MY_ID}_JELLYFIN_CONTAINER" "$container" "Jellyfin container name"
|
||||
_set_conf_var "${MY_ID}_JELLYFIN_URL" "http://localhost:${host_port:-8095}" "Jellyfin URL"
|
||||
;;
|
||||
esac
|
||||
|
||||
# Transcode path: first container with a /transcode mount wins
|
||||
if [[ -z "$transcode_dir" ]]; then
|
||||
transcode_dir=$(_docker_volume_host "$container" "/transcode")
|
||||
fi
|
||||
done
|
||||
|
||||
[[ -n "$transcode_dir" ]] && \
|
||||
_set_conf_var "${MY_ID}_TRANSCODE_SSD" "${transcode_dir%/}/" "Transcode SSD path"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Authelia container + config path ──────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
authelia_container=$(docker ps -a --format '{{.Names}}' 2>/dev/null | grep -im1 "authelia")
|
||||
if [[ -n "$authelia_container" ]]; then
|
||||
_set_conf_var "${MY_ID}_AUTHELIA_CONTAINER" "$authelia_container" "Authelia container"
|
||||
authelia_cfg_dir=$(_docker_volume_host "$authelia_container" "/config")
|
||||
[[ -n "$authelia_cfg_dir" ]] && \
|
||||
_set_conf_var "${MY_ID}_AUTHELIA_CONFIG" "${authelia_cfg_dir}/configuration.yml" "Authelia config path"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Network interface ─────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
nic=$(ip route show default 2>/dev/null | grep -oP '(?<=dev )\S+' | head -1)
|
||||
_set_conf_var "${MY_ID}_SYS_WATCHDOG_NIC" "$nic" "Default NIC"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Storage mode (internal NVMe/SSD vs USB flash boot) ───────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
boot_src=$(findmnt -n -o SOURCE /boot 2>/dev/null)
|
||||
if [[ -n "$boot_src" ]]; then
|
||||
boot_dev=$(lsblk -no PKNAME "$boot_src" 2>/dev/null | head -1)
|
||||
if [[ "$boot_dev" =~ ^nvme ]]; then
|
||||
storage_mode=true
|
||||
else
|
||||
boot_transport=$(cat "/sys/block/${boot_dev}/device/transport" 2>/dev/null)
|
||||
[[ "$boot_transport" == "usb" ]] && storage_mode=false || storage_mode=true
|
||||
fi
|
||||
_set_conf_var "${MY_ID}_STORAGE_MODE_INTERNAL" "$storage_mode" \
|
||||
"Storage mode (NVMe/SSD=true, USB=false)"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Master conf: host identity + Gitea container ─────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
if [[ -f "$MASTER_CONF" ]]; then
|
||||
echo ""
|
||||
echo " ── Master conf ──"
|
||||
_set_conf_var "$MY_ID" "$LOCAL_SERVER_NAME" "master.conf ${MY_ID} hostname" "$MASTER_CONF"
|
||||
|
||||
gitea_container=$(docker ps -a --format '{{.Names}}' 2>/dev/null | grep -im1 "gitea")
|
||||
[[ -n "$gitea_container" ]] && \
|
||||
_set_conf_var "GITEA_CONTAINER" "$gitea_container" "Gitea container" "$MASTER_CONF"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Summary + push ────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY Conf Populate Summary ━━━━━"
|
||||
echo " Updated: $UPDATED field(s)"
|
||||
echo " Skipped: $SKIPPED already set"
|
||||
echo " Conf: $CONF_FILE"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
if [[ "$UPDATED" -gt 0 ]] && [[ "$DRY_RUN" == false ]] && [[ "$NO_PUSH" == false ]]; then
|
||||
echo ""
|
||||
info "Pushing updated conf to partners..."
|
||||
bash "$SCRIPTS_ROOT/System_Essentials/conf_sync.sh" --push-only "${EXTRA_FLAGS[@]}" || true
|
||||
fi
|
||||
Executable
+469
@@ -0,0 +1,469 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================= conf_upgrade.sh ================================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Merges a new conf template into an existing user conf while preserving every
|
||||
# value the user has already set. Run manually when the conf schema changes
|
||||
# between versions — adds new keys, removes deprecated ones, and keeps the
|
||||
# structure of the new template exactly.
|
||||
#
|
||||
# Keys in template only → ADDED (placeholder/default — user fills in once)
|
||||
# Keys in target only → REMOVED (deprecated in new version)
|
||||
# Keys in both → KEPT (target's value always wins, template ignored)
|
||||
# Comments / blank lines → always from template (structure follows new version)
|
||||
#
|
||||
# Supports all conf variable patterns: simple scalars, indexed arrays, and
|
||||
# associative arrays (declare -A).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Validate both --template and --target exist
|
||||
# 2. Parse each into keys, handling scalars, indexed arrays and declare -A
|
||||
# 3. Classify every key: ADDED (template only) / REMOVED (target only) / KEPT (both)
|
||||
# 4. Emit the merged result — template structure, target values — to a temp file
|
||||
# staged in the target's own directory
|
||||
# 5. --dry-run stops here and prints the report
|
||||
# 6. --backup copies the current target to .bak
|
||||
# 7. Install by atomic rename over the target
|
||||
#
|
||||
# Called automatically by git_pull_execute.sh after every pull, for master.conf and this
|
||||
# host's own host*.conf. The host template is HOSTN_-prefixed and the caller substitutes
|
||||
# the real MY_ID before merging.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# The User's Value Always Wins
|
||||
# For any key present in both files, the target's value is kept and the template's is
|
||||
# discarded. The template supplies structure and new keys, never settings. This is what
|
||||
# makes the upgrade safe to run unattended after every single pull.
|
||||
#
|
||||
# Structure Follows the Template
|
||||
# Comments, ordering and blank lines come from the template, so an upgraded conf reads
|
||||
# like the current version rather than accumulating layers of old formatting.
|
||||
#
|
||||
# Standalone by Design — Do Not Add load_config.sh
|
||||
# This script sources nothing. It is the tool that repairs the conf that load_config.sh
|
||||
# depends on, so it has to work when that conf is broken, partial, or missing keys.
|
||||
# Sourcing load_config.sh here would make the repair tool fail in exactly the situation
|
||||
# it exists for. That is also why it uses plain echo instead of log()/error(), and why
|
||||
# there is no acquire_lock — common.sh is not available to it.
|
||||
#
|
||||
# Atomic Install, Never In-Place
|
||||
# The merged conf is renamed over the target, not copied into it. Every watchdog sources
|
||||
# load_config.sh on every run; a cp would truncate master.conf and write into it, and
|
||||
# anything reading during that window gets a partial conf with empty path variables.
|
||||
#
|
||||
# Concurrency Handled by Atomicity, Not a Lock
|
||||
# Two concurrent runs against the same target cannot corrupt it — each stages its own
|
||||
# temp file and the rename is atomic, so the last writer simply wins. Since the merge is
|
||||
# idempotent, that outcome is identical to running once.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# The Host Slot Must Be Declared, Never Inferred
|
||||
# host.conf.template ships HOSTN_ placeholders; a live host conf uses HOST1_ / HOST2_. Key
|
||||
# matching is literal, so merging the raw template against a real host conf classifies every
|
||||
# existing key as deprecated — KEPT 0 — and the install drops every credential in the file.
|
||||
# Verified on HOST1 2026-08-02: it would have removed HOST1_RADARR_API_KEY,
|
||||
# HOST1_EMBY_API_KEY, HOST1_NPM_PASS and 137 others.
|
||||
#
|
||||
# --host-slot HOST<n> is how a caller says which slot the template is being resolved for, and
|
||||
# it works for any slot — HOST1, HOST2, and whatever a third server would be. The slot is
|
||||
# still never inferred from the target and applied silently: the caller declares it, and the
|
||||
# target is read only to contradict a wrong answer. A declared slot that disagrees with the
|
||||
# target's own keys aborts, because substituting for the wrong slot destroys the file just as
|
||||
# thoroughly as not substituting at all. Without the flag, a template containing HOSTN is
|
||||
# refused exactly as before.
|
||||
#
|
||||
# Both cases are substituted. HOSTN_ covers the key prefixes, bare HOSTN appears in section
|
||||
# comments, and lowercase hostn is a real value — the hostn-appdata rsync profile keys. A
|
||||
# substitution handling only HOSTN_ leaves a conf carrying a profile named hostn-appdata that
|
||||
# nothing references.
|
||||
#
|
||||
# A Target Owning Two Slots Is Refused
|
||||
# A host conf describes exactly one server. Finding both HOST1_ and HOST2_ key definitions in
|
||||
# one target means it is not the file it claims to be, so the slot cross-check has nothing
|
||||
# trustworthy to compare against and the run aborts rather than picking one.
|
||||
#
|
||||
# Total Mismatch Is Refused
|
||||
# Keeping nothing from a populated conf is never a real upgrade; it means the two files do
|
||||
# not describe the same thing. KEPT 0 with a non-empty REMOVED list aborts. This is the
|
||||
# general net behind the HOSTN check — the failure mode is silent and total, so it fires in
|
||||
# --dry-run as well, putting the warning in the report itself.
|
||||
#
|
||||
# Root Required to Write
|
||||
# Installing over the target requires root. --dry-run deliberately does not, so the
|
||||
# change report can be previewed by anyone.
|
||||
#
|
||||
# Atomic Rename
|
||||
# The temp file is created in the target's own directory so the rename stays within one
|
||||
# filesystem. On Unraid /tmp is rootfs while the confs live on flash, and a cross-device
|
||||
# mv degrades to copy-then-unlink — precisely the torn write this avoids.
|
||||
#
|
||||
# Permissions Preserved
|
||||
# mktemp creates 0600; the target's existing mode and owner are copied onto the temp file
|
||||
# before it is installed, so a conf does not come back with different permissions.
|
||||
#
|
||||
# Temp File Cleanup
|
||||
# An EXIT trap removes the staged file on any early exit, and is cleared once the rename
|
||||
# has succeeded so the trap cannot delete the installed conf.
|
||||
#
|
||||
# Dry-run Mode
|
||||
# --dry-run prints the full change report (ADDED / REMOVED / KEPT) then exits
|
||||
# without writing anything. Always preview before applying to production confs.
|
||||
#
|
||||
# Backup Option
|
||||
# --backup writes a .bak copy of the target before overwriting. Use when
|
||||
# applying to a conf that has never been upgraded before.
|
||||
#
|
||||
# File Existence Guards
|
||||
# Both --template and --target are validated before any parsing begins.
|
||||
# Missing files abort immediately with a clear error.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# No conf vars. All inputs are CLI flags.
|
||||
#
|
||||
# --template <file> New version conf file (source of structure and defaults)
|
||||
# --target <file> Existing user conf (source of real values — always preserved)
|
||||
# --host-slot <HOSTn> Resolve HOSTN/hostn placeholders to this slot before merging.
|
||||
# Required for host.conf.template; meaningless for master.conf.template,
|
||||
# which has no placeholders. Must match the target's own slot.
|
||||
# --dry-run Show what would change without writing
|
||||
# --backup Write a .bak copy of target before modifying
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# conf_upgrade.sh --template Deployment/master.conf.template --target Configurations/master.conf --dry-run
|
||||
# Preview what would be added, removed, and kept — no changes written.
|
||||
#
|
||||
# conf_upgrade.sh --template Deployment/master.conf.template --target Configurations/master.conf --backup
|
||||
# Apply the upgrade, writing a .bak first.
|
||||
#
|
||||
# conf_upgrade.sh --template Deployment/master.conf.template --target Configurations/master.conf
|
||||
# Apply the upgrade in-place with no backup.
|
||||
#
|
||||
# conf_upgrade.sh --template Deployment/host.conf.template --target Configurations/host1.conf \
|
||||
# --host-slot HOST1 --dry-run
|
||||
# Preview a host conf upgrade. HOSTN/hostn are resolved to HOST1/host1 first. Swap in HOST2
|
||||
# and host2.conf for the other server — the template is the same file for every slot.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Arguments ────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
TEMPLATE=""
|
||||
TARGET=""
|
||||
DRY_RUN=false
|
||||
BACKUP=false
|
||||
HOST_SLOT=""
|
||||
_RESOLVED_TMPL="" # set only when --host-slot triggers a substitution; cleaned on exit
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--template) TEMPLATE="$2"; shift 2 ;;
|
||||
--target) TARGET="$2"; shift 2 ;;
|
||||
--host-slot) HOST_SLOT="$2"; shift 2 ;;
|
||||
--dry-run) DRY_RUN=true; shift ;;
|
||||
--backup) BACKUP=true; shift ;;
|
||||
*) echo "Unknown option: $1" >&2; exit 1 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
[[ -z "$TEMPLATE" ]] && { echo "Error: --template required" >&2; exit 1; }
|
||||
[[ -z "$TARGET" ]] && { echo "Error: --target required" >&2; exit 1; }
|
||||
[[ -f "$TEMPLATE" ]] || { echo "Error: template not found: $TEMPLATE" >&2; exit 1; }
|
||||
[[ -f "$TARGET" ]] || { echo "Error: target not found: $TARGET" >&2; exit 1; }
|
||||
|
||||
# ── Host slot resolution ─────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# host.conf.template ships HOSTN_ placeholders; a live host conf uses HOST1_ / HOST2_. Key
|
||||
# matching below is literal, so merging the raw template against a real host conf classifies
|
||||
# EVERY existing key as deprecated and every template key as new — KEPT 0, and the install
|
||||
# would drop every credential in the file.
|
||||
#
|
||||
# --host-slot is how a caller declares which slot the template is for. The slot is never
|
||||
# inferred from the target and silently applied: the caller states it, and the target is used
|
||||
# only to contradict a wrong answer. Substituting for the wrong slot is the same catastrophe as
|
||||
# not substituting at all, so a declared slot that disagrees with the target is refused.
|
||||
#
|
||||
# Both cases matter. HOSTN_ covers the 159 key prefixes; bare HOSTN appears in section comments,
|
||||
# and lowercase hostn is a real value — the hostn-appdata rsync profile keys. A substitution
|
||||
# that only handles HOSTN_ leaves a live conf with a profile named hostn-appdata that nothing
|
||||
# references, which is what the pull script did before this flag existed.
|
||||
_target_slots=$(grep -oE '^[[:space:]]*HOST[0-9]+_' "$TARGET" 2>/dev/null \
|
||||
| grep -oE 'HOST[0-9]+' | sort -u)
|
||||
_target_slot=$(echo "$_target_slots" | head -1)
|
||||
if [[ $(echo "$_target_slots" | grep -c .) -gt 1 ]]; then
|
||||
echo "Error: '$TARGET' defines keys for more than one host slot:" >&2
|
||||
echo " $(echo "$_target_slots" | tr '\n' ' ')" >&2
|
||||
echo " A host conf owns exactly one slot. Refusing rather than picking one." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ -n "$HOST_SLOT" ]]; then
|
||||
if ! [[ "$HOST_SLOT" =~ ^HOST[0-9]+$ ]]; then
|
||||
echo "Error: --host-slot must be HOST<n> (e.g. HOST1, HOST2) — got '$HOST_SLOT'" >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ -n "$_target_slot" && "$_target_slot" != "$HOST_SLOT" ]]; then
|
||||
echo "Error: --host-slot says $HOST_SLOT but '$TARGET' defines ${_target_slot}_ keys." >&2
|
||||
echo " Substituting for the wrong slot removes every key the target actually has," >&2
|
||||
echo " credentials included. Refusing." >&2
|
||||
exit 1
|
||||
fi
|
||||
if grep -q -e 'HOSTN' -e 'hostn' "$TEMPLATE" 2>/dev/null; then
|
||||
_lower=$(echo "$HOST_SLOT" | tr '[:upper:]' '[:lower:]')
|
||||
_RESOLVED_TMPL="$(mktemp)"
|
||||
trap '[[ -n "${_RESOLVED_TMPL:-}" ]] && rm -f "$_RESOLVED_TMPL"' EXIT
|
||||
sed -e "s/HOSTN/${HOST_SLOT}/g" -e "s/hostn/${_lower}/g" "$TEMPLATE" > "$_RESOLVED_TMPL"
|
||||
TEMPLATE="$_RESOLVED_TMPL"
|
||||
echo " Resolved HOSTN → ${HOST_SLOT} for this host"
|
||||
fi
|
||||
elif grep -q 'HOSTN' "$TEMPLATE" 2>/dev/null; then
|
||||
# No slot declared and the template is still generic — the original refusal, unchanged.
|
||||
if [[ -n "$_target_slot" ]]; then
|
||||
_lower=$(echo "$_target_slot" | tr '[:upper:]' '[:lower:]')
|
||||
echo "Error: template still contains HOSTN placeholders, but the target uses ${_target_slot}_." >&2
|
||||
echo " Merging as-is would classify all ${_target_slot}_ keys as deprecated and remove" >&2
|
||||
echo " them — including every credential. Declare the slot:" >&2
|
||||
echo "" >&2
|
||||
echo " $0 --template $TEMPLATE --target $TARGET --host-slot ${_target_slot} --dry-run" >&2
|
||||
echo "" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Block-parsing patterns ───────────────────────────────────────────────────────────────────
|
||||
# All three walkers below share these. A single-line array — KEY=(a b c) — both opens and
|
||||
# closes on one line. Matching only the opener leaves the walker inside a block it never
|
||||
# leaves, so every key until the next standalone ")" becomes invisible: skipped by
|
||||
# _collect_stats, dropped from HOST_MAP by _parse_target, and emitted from the template
|
||||
# instead of the user's conf by _write_merged. _ARR_ONELINE_RE is what stops that.
|
||||
|
||||
_ARR_DECL_RE='^[[:space:]]*declare[[:space:]]+-[a-zA-Z]+[[:space:]]+([A-Z0-9_]+)[[:space:]]*=\('
|
||||
_ARR_OPEN_RE='^[[:space:]]*([A-Z0-9_]+)[[:space:]]*=\('
|
||||
_ARR_CLOSE_RE='^[[:space:]]*\)[[:space:]]*(#.*)?$'
|
||||
# The [^#]* is deliberate: it forces the closing ")" to appear before any comment, so an
|
||||
# opener like FOO=( # see note (here) is not mistaken for a complete single-line array.
|
||||
_ARR_ONELINE_RE='=\([^#]*\)[[:space:]]*(#.*)?$'
|
||||
_SCALAR_RE='^[[:space:]]*([A-Z0-9_]+)[[:space:]]*='
|
||||
|
||||
# ── Parse target → KEY → full definition block ───────────────────────────────────────────────
|
||||
|
||||
declare -A HOST_MAP # KEY → complete definition line(s) from user's conf
|
||||
|
||||
_parse_target() {
|
||||
local in_block=false cur_key="" cur_block="" line
|
||||
|
||||
while IFS= read -r line || [[ -n "$line" ]]; do
|
||||
if [[ "$in_block" == true ]]; then
|
||||
cur_block+="$line"$'\n'
|
||||
# Closing ) — optional trailing whitespace and comment
|
||||
if [[ "$line" =~ $_ARR_CLOSE_RE ]]; then
|
||||
HOST_MAP["$cur_key"]="$cur_block"
|
||||
in_block=false; cur_key=""; cur_block=""
|
||||
fi
|
||||
else
|
||||
# declare -A KEY=( / KEY=(
|
||||
if [[ "$line" =~ $_ARR_DECL_RE ]] || [[ "$line" =~ $_ARR_OPEN_RE ]]; then
|
||||
cur_key="${BASH_REMATCH[1]}"
|
||||
# KEY=(a b c) — opens and closes on one line, never enter block mode
|
||||
if [[ "$line" =~ $_ARR_ONELINE_RE ]]; then
|
||||
HOST_MAP["$cur_key"]="$line"$'\n'
|
||||
cur_key=""
|
||||
else
|
||||
in_block=true; cur_block="$line"$'\n'
|
||||
fi
|
||||
# KEY=value (simple scalar)
|
||||
elif [[ "$line" =~ $_SCALAR_RE ]]; then
|
||||
HOST_MAP["${BASH_REMATCH[1]}"]="$line"$'\n'
|
||||
fi
|
||||
fi
|
||||
done < "$TARGET"
|
||||
}
|
||||
|
||||
# ── Walk template — collect stats (must run in current shell so arrays persist) ──────────────
|
||||
|
||||
declare -a ADDED=() KEPT=() REMOVED=()
|
||||
declare -A TMPL_SEEN=()
|
||||
|
||||
_collect_stats() {
|
||||
local in_block=false cur_key="" line
|
||||
|
||||
while IFS= read -r line || [[ -n "$line" ]]; do
|
||||
if [[ "$in_block" == true ]]; then
|
||||
if [[ "$line" =~ $_ARR_CLOSE_RE ]]; then
|
||||
in_block=false
|
||||
TMPL_SEEN["$cur_key"]=1
|
||||
if [[ -n "${HOST_MAP[$cur_key]+_}" ]]; then KEPT+=("$cur_key")
|
||||
else ADDED+=("$cur_key"); fi
|
||||
cur_key=""
|
||||
fi
|
||||
else
|
||||
if [[ "$line" =~ $_ARR_DECL_RE ]] || [[ "$line" =~ $_ARR_OPEN_RE ]]; then
|
||||
cur_key="${BASH_REMATCH[1]}"
|
||||
if [[ "$line" =~ $_ARR_ONELINE_RE ]]; then
|
||||
TMPL_SEEN["$cur_key"]=1
|
||||
if [[ -n "${HOST_MAP[$cur_key]+_}" ]]; then KEPT+=("$cur_key")
|
||||
else ADDED+=("$cur_key"); fi
|
||||
cur_key=""
|
||||
else
|
||||
in_block=true
|
||||
fi
|
||||
elif [[ "$line" =~ $_SCALAR_RE ]]; then
|
||||
local k="${BASH_REMATCH[1]}"
|
||||
TMPL_SEEN["$k"]=1
|
||||
if [[ -n "${HOST_MAP[$k]+_}" ]]; then KEPT+=("$k")
|
||||
else ADDED+=("$k"); fi
|
||||
fi
|
||||
fi
|
||||
done < "$TEMPLATE"
|
||||
|
||||
for key in "${!HOST_MAP[@]}"; do
|
||||
[[ -z "${TMPL_SEEN[$key]+_}" ]] && REMOVED+=("$key")
|
||||
done
|
||||
}
|
||||
|
||||
# ── Walk template — write merged output ──────────────────────────────────────────────────────
|
||||
# Runs in a subshell (stdout redirected) — array mutations are intentionally discarded here.
|
||||
|
||||
_write_merged() {
|
||||
local in_block=false cur_key="" cur_block="" line
|
||||
|
||||
while IFS= read -r line || [[ -n "$line" ]]; do
|
||||
if [[ "$in_block" == true ]]; then
|
||||
cur_block+="$line"$'\n'
|
||||
if [[ "$line" =~ $_ARR_CLOSE_RE ]]; then
|
||||
in_block=false
|
||||
if [[ -n "${HOST_MAP[$cur_key]+_}" ]]; then printf '%s' "${HOST_MAP[$cur_key]}"
|
||||
else printf '%s' "$cur_block"; fi
|
||||
cur_key=""; cur_block=""
|
||||
fi
|
||||
else
|
||||
if [[ "$line" =~ $_ARR_DECL_RE ]] || [[ "$line" =~ $_ARR_OPEN_RE ]]; then
|
||||
cur_key="${BASH_REMATCH[1]}"
|
||||
if [[ "$line" =~ $_ARR_ONELINE_RE ]]; then
|
||||
if [[ -n "${HOST_MAP[$cur_key]+_}" ]]; then printf '%s' "${HOST_MAP[$cur_key]}"
|
||||
else printf '%s\n' "$line"; fi
|
||||
cur_key=""
|
||||
else
|
||||
in_block=true; cur_block="$line"$'\n'
|
||||
fi
|
||||
elif [[ "$line" =~ $_SCALAR_RE ]]; then
|
||||
local k="${BASH_REMATCH[1]}"
|
||||
if [[ -n "${HOST_MAP[$k]+_}" ]]; then printf '%s' "${HOST_MAP[$k]}"
|
||||
else printf '%s\n' "$line"; fi
|
||||
else
|
||||
printf '%s\n' "$line"
|
||||
fi
|
||||
fi
|
||||
done < "$TEMPLATE"
|
||||
}
|
||||
|
||||
# ── Run ───────────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
_parse_target
|
||||
_collect_stats
|
||||
|
||||
# ── Report ────────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
TARGET_NAME="$(basename "$TARGET")"
|
||||
echo ""
|
||||
echo "── conf_upgrade: $TARGET_NAME ──────────────────────────────────────────"
|
||||
|
||||
if [[ ${#ADDED[@]} -gt 0 ]]; then
|
||||
echo " ADDED (new — fill in your values where needed):"
|
||||
for k in "${ADDED[@]}"; do echo " + $k"; done
|
||||
fi
|
||||
|
||||
if [[ ${#REMOVED[@]} -gt 0 ]]; then
|
||||
echo " REMOVED (deprecated — no longer in this version):"
|
||||
for k in "${REMOVED[@]}"; do echo " - $k"; done
|
||||
fi
|
||||
|
||||
echo " KEPT ${#KEPT[@]} existing vars — your values preserved"
|
||||
|
||||
# ── Guard: total mismatch ────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# Keeping nothing from a populated conf is never a real upgrade — it is the signature of the
|
||||
# two files not describing the same thing (wrong template, wrong target, unsubstituted
|
||||
# placeholders). The HOSTN check above catches the known cause; this catches the rest, because
|
||||
# the failure mode is silent and total: every value in the file is replaced by a default.
|
||||
# Deliberately fires in --dry-run too, so the report itself carries the warning.
|
||||
if [[ ${#KEPT[@]} -eq 0 && ${#REMOVED[@]} -gt 0 ]]; then
|
||||
echo "────────────────────────────────────────────────────────────────────────"
|
||||
echo ""
|
||||
echo "Error: refusing — this would keep NOTHING and remove all ${#REMOVED[@]} existing keys." >&2
|
||||
echo " A genuine upgrade preserves values; keeping zero means the template and the" >&2
|
||||
echo " target do not describe the same conf. Check that --template matches --target" >&2
|
||||
echo " and that any HOSTN placeholders were substituted for this host's slot." >&2
|
||||
echo "" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ${#ADDED[@]} -eq 0 && ${#REMOVED[@]} -eq 0 ]]; then
|
||||
echo " Already up to date — no changes needed."
|
||||
echo "────────────────────────────────────────────────────────────────────────"
|
||||
echo ""
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "────────────────────────────────────────────────────────────────────────"
|
||||
echo ""
|
||||
|
||||
# ── Apply ─────────────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo "(dry-run — no changes written)"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Checked here rather than at the top: --dry-run is a read-only report and is useful to
|
||||
# anyone, but installing over a conf under /boot needs root. Plain echo because this script
|
||||
# deliberately does not source common.sh — see the header.
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
echo "ERROR: writing '$TARGET' requires root (use --dry-run to preview as any user)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Staged beside the target, not in /tmp. mv is only atomic within one filesystem, and on
|
||||
# Unraid /tmp is rootfs while the confs live on flash — a cross-device mv silently degrades
|
||||
# to copy-then-unlink, which is exactly the torn write this is meant to prevent.
|
||||
TMPOUT="$(mktemp "${TARGET}.XXXXXX")"
|
||||
trap 'rm -f "$TMPOUT"; [[ -n "${_RESOLVED_TMPL:-}" ]] && rm -f "$_RESOLVED_TMPL"' EXIT
|
||||
|
||||
# mktemp creates 0600; carry the target's existing mode/owner across so the installed conf
|
||||
# does not come back with different permissions than it went in with.
|
||||
chmod --reference="$TARGET" "$TMPOUT" 2>/dev/null || true
|
||||
chown --reference="$TARGET" "$TMPOUT" 2>/dev/null || true
|
||||
|
||||
_write_merged > "$TMPOUT"
|
||||
|
||||
if [[ "$BACKUP" == true ]]; then
|
||||
cp -a "$TARGET" "${TARGET}.bak"
|
||||
echo "Backup: ${TARGET}.bak"
|
||||
fi
|
||||
|
||||
# Atomic install. cp would truncate the live conf and write into it, leaving a window where
|
||||
# anything sourcing load_config.sh reads a half-written master.conf — every watchdog does
|
||||
# that constantly. A rename swaps the inode: readers get the old file or the new one.
|
||||
mv -f "$TMPOUT" "$TARGET"
|
||||
trap - EXIT
|
||||
[[ -n "$_RESOLVED_TMPL" ]] && rm -f "$_RESOLVED_TMPL"
|
||||
echo "Updated: $TARGET"
|
||||
@@ -0,0 +1,648 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================== HOSTN CONFIGURATION — (hostname) ==================================
|
||||
# ==============================================================================================
|
||||
# HOSTN-specific variables — credentials, container names, share paths, failover lists.
|
||||
# Sourced after master.conf — values here extend shared profile arrays and add HOSTN-specific
|
||||
# identity, credentials, and container configuration.
|
||||
#
|
||||
# Sparse checkout (git) ensures other hosts never receive this file.
|
||||
#
|
||||
# DO NOT put shared config here — thresholds, toggles, profiles belong in master.conf.
|
||||
# DO NOT put other hosts' variables here — they belong in their own host*.conf files.
|
||||
#
|
||||
# ── HOW TO USE THIS TEMPLATE ──────────────────────────────────────────────────────────────────
|
||||
# This file was generated by the Varaverk first-run wizard.
|
||||
# Fill in the sections that apply to your setup — leave unused sections empty.
|
||||
# All scripts self-guard against empty values — safe to leave sections blank until needed.
|
||||
#
|
||||
# ── INDEX ─────────────────────────────────────────────────────────────────────────────────────
|
||||
#
|
||||
# ── IDENTITY & CONNECTIVITY ────────────────────────────────────────────────────────────────
|
||||
# IDENTITY hostname, SSH key, Unraid API key
|
||||
# EMBY container name, URL, API key
|
||||
# JELLYFIN container name, URL, API key
|
||||
# GITEA API token for SSH key registration
|
||||
# NOTIFICATIONS Discord webhook
|
||||
#
|
||||
# ── PARTNERSHIP ────────────────────────────────────────────────────────────────────────────
|
||||
# PARTNERSHIP auth containers, backup paths, emby provisioning
|
||||
#
|
||||
# ── RSYNC ──────────────────────────────────────────────────────────────────────────────────
|
||||
# DAILY SYNC SHARES media shares this host owns and pushes
|
||||
# PERSONAL SHARES private encrypted shares for offsite backup
|
||||
# WEEKLY SYNC SHARES appdata shares synced weekly
|
||||
# INTERMEDIATE SYNC mid-day appdata propagation
|
||||
# CRITICAL SYNC SHARES appdata shares synced every 30 minutes
|
||||
# BACKUP VERIFY shares for checksum verification against remote
|
||||
# HOSTN RSYNC PROFILE host-specific appdata sync profile
|
||||
#
|
||||
# ── FALLBACK ───────────────────────────────────────────────────────────────────────────────
|
||||
# DDNS DDNS containers managed by this host
|
||||
# INTERNET LOSS containers stopped when internet is lost
|
||||
# FALLBACK TIERS what this host wants the partner to run when this host is down
|
||||
# TIER DELAYS how long this host must be down before each tier activates
|
||||
# RSYNC WRITEBACK appdata synced back on handback
|
||||
#
|
||||
# ── DOCKER ─────────────────────────────────────────────────────────────────────────────────
|
||||
# DOCKER DAILY RESTART containers restarted daily
|
||||
# DOCKER WEEKLY RESTART containers restarted weekly
|
||||
# DOCKER WATCHDOG memory limits, health URLs, required containers
|
||||
# NETWORK WATCHDOG DDNS domain, NPM URL for connectivity checks
|
||||
# DOCKER NETWORK CONNECT networks and containers for array start
|
||||
#
|
||||
# ── MEDIA ──────────────────────────────────────────────────────────────────────────────────
|
||||
# MEDIA PERMISSIONS share list for permissions script
|
||||
# MEDIA CLEANER folder lists for media_cleaner.sh
|
||||
#
|
||||
# ── ARR STACK ──────────────────────────────────────────────────────────────────────────────
|
||||
# DOWNLOADERS slskd, SABnzbd, qBittorrent credentials and URLs
|
||||
# LIDARR / SONARR / RADARR URL, API key, path map
|
||||
# ARR RECOVERY per-arr recovery toggles
|
||||
#
|
||||
# ── TRANSCODES ─────────────────────────────────────────────────────────────────────────────
|
||||
# TRANSCODES ramdisk size, thresholds, SSD path, server array
|
||||
#
|
||||
# ── MONITORS ───────────────────────────────────────────────────────────────────────────────
|
||||
# CERTIFICATE MONITOR domains checked for SSL expiry
|
||||
# SMART HEALTH drives to skip in SMART monitoring
|
||||
# ZFS REPORT pools to exclude from ZFS health report
|
||||
# PCIe AER QUIET dead PCIe devices removed at array start to stop AER log spam
|
||||
#
|
||||
# ── RESOURCE MANAGER ───────────────────────────────────────────────────────────────────────
|
||||
# RESOURCE MANAGER containers paused/stopped under memory pressure
|
||||
#
|
||||
# ── SYSTEM WATCHDOG ────────────────────────────────────────────────────────────────────────
|
||||
# SYSTEM WATCHDOG per-host check toggles and NIC configuration
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
# ==============================================================================================
|
||||
# ── STORAGE MODE ──────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Storage mode ━━━
|
||||
# Controls where Varaverk stores scripts, conf, and state files.
|
||||
# true = internal NVMe/SSD — /boot/config/plugins/varaverk (write-safe, git-direct)
|
||||
# false = USB flash boot — /mnt/user/appdata/Varaverk (preserves flash lifetime)
|
||||
# Auto-detected from boot device transport on first setup.
|
||||
# To change: Settings → Storage → Migrate.
|
||||
HOSTN_STORAGE_MODE_INTERNAL=false
|
||||
|
||||
# ==============================================================================================
|
||||
# ── IDENTITY & CONNECTIVITY ───────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Identity ━━━
|
||||
# HOSTN hostname lives in master.conf (not a credential — safe for all servers).
|
||||
# SSH key used for all server-to-server operations — rsync, fallback, conf sync.
|
||||
# Convention: /root/.ssh/<hostname-lowercase-no-unraid-prefix>_rsync_automation
|
||||
# Must be in /root/.ssh/ and authorised in the partner's /root/.ssh/authorized_keys.
|
||||
# Run Partnership/ssh_setup.sh to generate the key and copy it to the partner.
|
||||
HOSTN_SSH_KEY="" # e.g. /root/.ssh/myserver_rsync_automation
|
||||
HOSTN_STORAGE_PATH="/mnt/user"
|
||||
HOSTN_OWNER="" # short identifier for this server (e.g. myserver)
|
||||
HOSTN_OWNER_EMAIL=""
|
||||
|
||||
# ━━━ Unraid API ━━━
|
||||
# Used by the Varaverk plugin to query this server's Unraid GraphQL API.
|
||||
# Generate in Unraid: Settings → Management Access → API Keys → + New Key
|
||||
HOSTN_UNRAID_API_KEY=""
|
||||
|
||||
# ━━━ Emby ━━━
|
||||
HOSTN_EMBY_CONTAINER="Emby"
|
||||
HOSTN_EMBY_URL="http://localhost:8096"
|
||||
HOSTN_EMBY_API_KEY="" # Emby Dashboard → API Keys → + New Key
|
||||
HOSTN_EMBY_PUBLIC_URL="" # e.g. https://media.example.com/emby — browser-reachable base, used
|
||||
# to build image URLs that render in the WebGUI. Deliberately separate
|
||||
# from HOSTN_EMBY_URL: that one is for server-side API calls and is
|
||||
# usually localhost, which resolves to the wrong machine in a browser.
|
||||
# Empty = features that need an image quietly go without one.
|
||||
|
||||
# ━━━ Jellyfin ━━━
|
||||
HOSTN_JELLYFIN_CONTAINER="Jellyfin"
|
||||
HOSTN_JELLYFIN_URL="http://localhost:8095"
|
||||
HOSTN_JELLYFIN_API_KEY="" # Jellyfin Dashboard → Administration → API Keys
|
||||
|
||||
# ━━━ Gitea ━━━
|
||||
# Personal access token for gitea_ssh_setup.sh.
|
||||
# Create in Gitea: Settings → Applications → Generate Token → scope: write:user
|
||||
HOSTN_GITEA_API_TOKEN=""
|
||||
|
||||
# ━━━ Bug Reports ━━━
|
||||
# Only used when BUG_REPORT_LOCAL_ENABLED=true in master.conf. Reports then go to this Gitea
|
||||
# instead of GitHub — and stay there, so they do not reach the Varaverk maintainer.
|
||||
#
|
||||
# Reached locally or over Tailscale, so the token never crosses the public proxy and no Authelia
|
||||
# bypass is needed. It is a credential and lives here rather than master.conf for that reason;
|
||||
# it is never shipped, and the settings UI masks it.
|
||||
HOSTN_BUG_REPORT_URL="" # e.g. http://gitea:3000 or the tailnet name
|
||||
HOSTN_BUG_REPORT_REPO="" # owner/repo
|
||||
HOSTN_BUG_REPORT_TOKEN="" # Gitea API token with issue-write on that repo
|
||||
|
||||
# ━━━ Notifications ━━━
|
||||
# Discord webhook — leave blank to disable.
|
||||
HOSTN_DISCORD_WEBHOOK=""
|
||||
|
||||
# ==============================================================================================
|
||||
# ── PARTNERSHIP ───────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Auth containers reconfigured on onboard/offboard.
|
||||
# Format: "ContainerName|WebUIPort"
|
||||
HOSTN_PARTNERSHIP_AUTH_WEBUIS=(
|
||||
# "NginxProxyManager|81"
|
||||
# "Authelia|9091"
|
||||
)
|
||||
|
||||
# Shares rsynced to the mirror during onboard Step 1e, BEFORE the auth containers are created.
|
||||
# This is the only rsync an onboard performs — media is never seeded here.
|
||||
# Profile is inferred from the directory basename, so Critical-Data resolves to critical-data:
|
||||
# a clean copy with the auth containers stopped on both sides. Do not point this at a share
|
||||
# whose profile keeps databases running; a dirty copy of MariaDB or Redis is worse than none,
|
||||
# because the container starts, reads Up, and restarts a dead database behind it.
|
||||
HOSTN_PARTNERSHIP_PROVISION_SHARES=(
|
||||
"/mnt/user/appdata-Fallback/Critical-Data" # critical-data profile — the auth stack
|
||||
)
|
||||
|
||||
# Containers that belong in "<PartnerShort>-Fallback" on the mirror rather than in a mirrored
|
||||
# copy of this host's folder layout — the ones that exist there only to cover this host going
|
||||
# dark. Everything else the onboard deploys is filed onto the same shelf it occupies here
|
||||
# (Arrs Stack, Networking, Databases…), because it runs on the mirror continuously.
|
||||
# Empty is the normal state: leave it empty unless a container is genuinely failover-only.
|
||||
HOSTN_PARTNERSHIP_FALLBACK_ONLY=()
|
||||
|
||||
# XML templates pushed to mirror during onboard — auth stack.
|
||||
# Dependencies (databases) must come before apps that depend on them.
|
||||
HOSTN_PARTNERSHIP_AUTH_STACK=(
|
||||
# "my-Authelia.xml"
|
||||
# "my-NginxProxyManager.xml"
|
||||
)
|
||||
|
||||
# XML templates pushed to mirror during onboard — arr stack.
|
||||
HOSTN_PARTNERSHIP_ARR_STACK=(
|
||||
# "my-Sonarr.xml"
|
||||
# "my-Radarr.xml"
|
||||
)
|
||||
|
||||
# XML templates pushed to mirror during onboard — other services (not auth, not arr).
|
||||
HOSTN_PARTNERSHIP_SERVICES_STACK=(
|
||||
# "my-Lidarr.xml"
|
||||
)
|
||||
|
||||
# Paths the partner should collect during the grace window after offboard.
|
||||
HOSTN_PARTNERSHIP_MIRROR_BACKUPS=(
|
||||
# "/mnt/user/appdata-Fallback/Partner-Emby"
|
||||
)
|
||||
|
||||
# Containers parked on this server when partnership is active.
|
||||
HOSTN_PARTNERSHIP_OWN_CONTAINERS=(
|
||||
# "Emby"
|
||||
)
|
||||
|
||||
# Containers stopped on THIS server before deploying the mirror's stack on onboard.
|
||||
HOSTN_PARTNERSHIP_REPLACE_CONTAINERS=(
|
||||
)
|
||||
|
||||
# Arr containers stopped on this server when mirror's arr stack is deployed.
|
||||
HOSTN_PARTNERSHIP_ARR_REPLACE_CONTAINERS=(
|
||||
)
|
||||
|
||||
# Emby admin provisioning — owner controls whether Emby is shared.
|
||||
HOSTN_PARTNERSHIP_PROVISION_EMBY_ADMIN=false
|
||||
HOSTN_PARTNERSHIP_EMBY_PORT=8096
|
||||
HOSTN_PARTNERSHIP_EMBY_ADMIN_USER=""
|
||||
HOSTN_PARTNERSHIP_EMBY_ADMIN_PASS=""
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RSYNC ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Daily Sync Shares ━━━
|
||||
# Media shares this host pushes to all other nodes every night.
|
||||
# Uses DEFAULT_RSYNC_OPTS from master.conf — no profile needed.
|
||||
HOSTN_DAILY_SYNC_SHARES=(
|
||||
# /mnt/user/Movies
|
||||
# /mnt/user/Tv_Shows
|
||||
# /mnt/user/Music
|
||||
)
|
||||
|
||||
# ━━━ Personal Shares ━━━
|
||||
# Private encrypted shares synced for offsite backup, independent of media shares.
|
||||
HOSTN_PERSONAL_SHARES=(
|
||||
# /mnt/user/Personal # e.g. ZFS-encrypted dataset
|
||||
)
|
||||
|
||||
# ━━━ Weekly Sync Shares ━━━
|
||||
# Appdata shares synced during the weekly maintenance window.
|
||||
# Profiles (emby, critical-data) drive container stops — define in master.conf.
|
||||
HOSTN_WEEKLY_SYNC_SHARES=(
|
||||
# "/mnt/user/Media_Server/Emby" # emby profile
|
||||
# "/mnt/user/appdata-Fallback/Critical-Data" # critical-data profile
|
||||
)
|
||||
|
||||
# ━━━ Monthly Sync Shares ━━━
|
||||
# Shares synced by monthly_maintenance.sh. Add here when ready.
|
||||
HOSTN_MONTHLY_SYNC_SHARES=(
|
||||
# Add shares here
|
||||
)
|
||||
|
||||
# ━━━ Intermediate Sync Shares ━━━
|
||||
# Shares synced every 4 hours. Leave empty to skip mid-day rsync.
|
||||
HOSTN_INTERMEDIATE_SYNC_SHARES=(
|
||||
# Add shares here to enable mid-day rsync
|
||||
)
|
||||
|
||||
# ━━━ Critical Sync Shares ━━━
|
||||
# Appdata shares synced every 30 minutes.
|
||||
# Format: "/path/to/share" or "/path/to/share|profile-name"
|
||||
HOSTN_CRITICAL_SYNC_SHARES=(
|
||||
# "/mnt/user/appdata-Fallback/Critical-Data|critical-fallback"
|
||||
)
|
||||
|
||||
# ━━━ Backup Verify ━━━
|
||||
# Leave empty to use HOSTN_DAILY_SYNC_SHARES automatically.
|
||||
HOSTN_BACKUP_VERIFY_SHARES=(
|
||||
# leave empty to use HOSTN_DAILY_SYNC_SHARES automatically
|
||||
)
|
||||
|
||||
# ━━━ HOSTN Rsync Profile — hostn-appdata ━━━
|
||||
# Host-specific appdata sync profile.
|
||||
# Retry/sleep/container-delay omitted — this profile uses the master.conf defaults for all three.
|
||||
PROFILE_RSYNC_OPTS[hostn-appdata]="-av --info=progress2 --bwlimit=${PROFILE_BW_LIMIT[hostn-appdata]:-8000}"
|
||||
PROFILE_BW_LIMIT[hostn-appdata]=8000
|
||||
PROFILE_CRITICAL_CONTAINER_NAMES[hostn-appdata]=""
|
||||
PROFILE_DELAYED_CONTAINERS[hostn-appdata]=""
|
||||
PROFILE_EXCLUDE_DIRS[hostn-appdata]="logs *.tmp"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FALLBACK ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ DDNS ━━━
|
||||
# DDNS containers this host manages.
|
||||
HOSTN_DDNS_CONTAINERS=(
|
||||
# "MyServer.com"
|
||||
)
|
||||
|
||||
# ━━━ Internet Loss ━━━
|
||||
# Containers stopped immediately when internet is lost.
|
||||
FALLBACK_HOSTN_STOP_ON_NO_NET=(
|
||||
# "MyServer.com"
|
||||
)
|
||||
|
||||
# ━━━ Fallback Tiers — What HOSTN Wants Covered When Down ━━━
|
||||
# Containers the partner starts for HOSTN when HOSTN goes down.
|
||||
# Tier 1 is always immediate. Higher tiers activate after HOSTN_TIER*_DELAY minutes.
|
||||
FALLBACK_HOSTN_TIER1=(
|
||||
# "MyServer-DDNS-Container"
|
||||
)
|
||||
|
||||
FALLBACK_HOSTN_TIER2=(
|
||||
# "container-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOSTN_TIER3=(
|
||||
# "container-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOSTN_TIER4=(
|
||||
# "container-placeholder"
|
||||
)
|
||||
|
||||
# ━━━ Tier Delays — HOSTN Outage Timers ━━━
|
||||
# How long HOSTN must be down before each tier activates on the partner.
|
||||
HOSTN_TIER2_DELAY=240 # 4 hours
|
||||
HOSTN_TIER3_DELAY=720 # 12 hours
|
||||
HOSTN_TIER4_DELAY=1440 # 24 hours
|
||||
|
||||
# ━━━ Rsync Writeback ━━━
|
||||
HOSTN_TIER1_WRITEBACK_DELAY=60 # minimum outage minutes before Tier 1 writeback runs
|
||||
|
||||
FALLBACK_HOSTN_WRITEBACK_TIER1=(
|
||||
# "/mnt/user/Media_Server/Emby"
|
||||
)
|
||||
|
||||
FALLBACK_HOSTN_WRITEBACK_TIER2=(
|
||||
# "/mnt/user/appdata-Fallback/Important-Data"
|
||||
)
|
||||
|
||||
FALLBACK_HOSTN_WRITEBACK_TIER3=(
|
||||
# "location-placeholder"
|
||||
)
|
||||
|
||||
FALLBACK_HOSTN_WRITEBACK_TIER4=(
|
||||
# "/mnt/user/appdata-Fallback/Arrs_Stack"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── DOCKER ────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Docker Daily Restart ━━━
|
||||
HOSTN_DAILY_RESTART_CONTAINERS=(
|
||||
# "NginxProxyManager"
|
||||
# "Authelia"
|
||||
)
|
||||
|
||||
# ━━━ Docker Weekly Restart ━━━
|
||||
HOSTN_WEEKLY_RESTART_CONTAINERS=(
|
||||
# "NextCloud"
|
||||
# "AdGuard-Home"
|
||||
)
|
||||
|
||||
# ━━━ Docker Watchdog ━━━
|
||||
|
||||
# Memory hard limits in MB — immediate restart if exceeded.
|
||||
# 20GB=20480 16GB=16384 8GB=8192 4GB=4096 2GB=2048 1GB=1024
|
||||
declare -A HOSTN_WATCHDOG_CONTAINERS=(
|
||||
# ["Emby"]=18432
|
||||
)
|
||||
|
||||
# HTTP health check URLs — checked every cycle.
|
||||
declare -A HOSTN_WATCHDOG_CONTAINER_URLS=(
|
||||
# ["Emby"]="http://localhost:8096"
|
||||
# ["Jellyfin"]="http://localhost:8095"
|
||||
# ["NginxProxyManager"]="http://localhost:7818"
|
||||
# ["Authelia"]="http://localhost:9091/api/health"
|
||||
# ["Authelia-Secondary"]="http://localhost:9092/api/health"
|
||||
# ["Lldap"]="http://localhost:17170"
|
||||
)
|
||||
|
||||
# API-level health checks — catches HTTP-200-but-internally-frozen containers (DB lock,
|
||||
# deadlocked thread, etc.) that a basic HTTP check above would miss. Pick an endpoint that
|
||||
# forces a real DB round-trip — a lightweight status endpoint may stay 200 even while the
|
||||
# rest of the app is locked up. Format: ["ContainerName"]="url|APIKey"
|
||||
declare -A HOSTN_WATCHDOG_CONTAINER_API_CHECKS=(
|
||||
# ["Emby"]="${HOSTN_EMBY_URL}/Users|${HOSTN_EMBY_API_KEY}"
|
||||
# ["Jellyfin"]="${HOSTN_JELLYFIN_URL}/Users|${HOSTN_JELLYFIN_API_KEY}"
|
||||
)
|
||||
|
||||
# Required containers — must always be running.
|
||||
HOSTN_WATCHDOG_REQUIRED_CONTAINERS=(
|
||||
# "NginxProxyManager"
|
||||
# "Authelia"
|
||||
)
|
||||
|
||||
# Containers to skip in Tier 2 global scan.
|
||||
HOSTN_WATCHDOG_SCAN_IGNORE=(
|
||||
# "my-occasional-container"
|
||||
)
|
||||
|
||||
# Dependency ordering — skip restarting a container if its dependency is also down.
|
||||
declare -A HOSTN_WATCHDOG_DEPENDENCIES=(
|
||||
# ["Authelia"]="Mariadb Redis-Authelia"
|
||||
)
|
||||
|
||||
# Per-container appdata growth suppress ceilings in MB.
|
||||
declare -A HOSTN_WATCHDOG_APPDATA_SIZES=(
|
||||
# ["Tdarr"]="25600"
|
||||
)
|
||||
|
||||
# ━━━ Network Watchdog ━━━
|
||||
HOSTN_NETWORK_WATCHDOG_DDNS_DOMAIN="" # e.g. myserver.com
|
||||
HOSTN_NETWORK_WATCHDOG_DDNS_CONTAINER="" # e.g. MyServer.com
|
||||
HOSTN_NETWORK_WATCHDOG_NPM_URL="" # e.g. https://myserver.com
|
||||
|
||||
# ━━━ Docker Network Connect ━━━
|
||||
HOSTN_NETWORK_CONNECT_CONTAINERS=(
|
||||
# "memcached"
|
||||
)
|
||||
|
||||
HOSTN_NETWORK_CONNECT_NETWORKS=(
|
||||
# "high-availability"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MEDIA ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Media Permissions ━━━
|
||||
HOSTN_MEDIA_PERMISSION_SHARES=(
|
||||
# /mnt/user/Movies
|
||||
# /mnt/user/Tv_Shows
|
||||
# /mnt/user/Music
|
||||
# /mnt/user/Downloads
|
||||
)
|
||||
|
||||
# ━━━ Media Cleaner ━━━
|
||||
HOSTN_ANIME_CLEAN_FOLDERS=(
|
||||
# /mnt/user/Anime_Movies
|
||||
# /mnt/user/Anime_Shows
|
||||
)
|
||||
|
||||
HOSTN_MEDIA_CLEAN_FOLDERS=(
|
||||
# /mnt/user/Movies
|
||||
# /mnt/user/Tv_Shows
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── ARR STACK ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Downloaders ━━━
|
||||
HOSTN_SLSKD_URL="http://localhost:8980"
|
||||
HOSTN_SLSKD_API_KEY=""
|
||||
HOSTN_SLSKD_FAILED_IMPORTS_DIR=""
|
||||
|
||||
HOSTN_SABNZBD_URL="http://localhost:8180"
|
||||
HOSTN_SABNZBD_API_KEY=""
|
||||
|
||||
HOSTN_QBIT_URL="http://localhost:8080"
|
||||
HOSTN_QBIT_USERNAME="admin"
|
||||
HOSTN_QBIT_PASSWORD=""
|
||||
|
||||
# ━━━ Lidarr ━━━
|
||||
HOSTN_LIDARR_URL="http://localhost:8686"
|
||||
HOSTN_LIDARR_API_KEY=""
|
||||
HOSTN_LIDARR_MUSIC_ROOT="/mnt/user/Music"
|
||||
HOSTN_FANART_API_KEY=""
|
||||
HOSTN_LASTFM_API_KEY=""
|
||||
|
||||
declare -A HOSTN_LIDARR_PATH_MAP=(
|
||||
# ["/music"]="/mnt/user/Music"
|
||||
)
|
||||
|
||||
# ━━━ Sonarr ━━━
|
||||
HOSTN_SONARR_URL="http://localhost:8989"
|
||||
HOSTN_SONARR_API_KEY=""
|
||||
HOSTN_SONARR_TV_ROOT="/mnt/user/Tv_Shows"
|
||||
HOSTN_SONARR_GENERAL_ROOT="" # rootFolderPath literal for the general root (e.g. "/tv") — target for reverse-kids-leak moves; leave blank to disable
|
||||
HOSTN_SONARR_KIDS_ROOT="" # rootFolderPath literal, as reported by Sonarr API — leave blank if no dedicated kids root
|
||||
HOSTN_SONARR_ANIME_ROOT="" # rootFolderPath literal, as reported by Sonarr API — leave blank if no dedicated anime root
|
||||
HOSTN_SONARR_DOWNLOAD_DIR="" # host path of the completed-downloads folder Sonarr imports from (e.g. "/mnt/cache/Temp_Storage/SABnzbd/Completed/Tv Shows") — blank disables the download orphan cleaner for Sonarr
|
||||
HOSTN_SONARR_DOWNLOAD_CONTAINER_DIR="" # same folder as Sonarr's container sees it (e.g. "/downloads/Completed/Tv Shows") — needed to trigger import scans on held folders
|
||||
|
||||
declare -A HOSTN_SONARR_PATH_MAP=(
|
||||
# ["/tv"]="/mnt/user/Tv_Shows"
|
||||
)
|
||||
|
||||
# ━━━ Radarr ━━━
|
||||
HOSTN_RADARR_URL="http://localhost:7878"
|
||||
HOSTN_RADARR_API_KEY=""
|
||||
HOSTN_TMDB_API_KEY=""
|
||||
HOSTN_RADARR_MOVIES_ROOT="/mnt/user/Movies"
|
||||
HOSTN_RADARR_GENERAL_ROOT="" # rootFolderPath literal for the general root (e.g. "/movies") — target for reverse-kids-leak moves; leave blank to disable
|
||||
HOSTN_RADARR_KIDS_ROOT="" # rootFolderPath literal, as reported by Radarr API — leave blank if no dedicated kids root
|
||||
HOSTN_RADARR_ANIME_ROOT="" # rootFolderPath literal, as reported by Radarr API — leave blank if no dedicated anime root
|
||||
HOSTN_RADARR_DOWNLOAD_DIR="" # host path of the completed-downloads folder Radarr imports from (e.g. "/mnt/cache/Temp_Storage/SABnzbd/Completed/Movies") — blank disables the download orphan cleaner for Radarr
|
||||
HOSTN_RADARR_DOWNLOAD_CONTAINER_DIR="" # same folder as Radarr's container sees it (e.g. "/downloads/Completed/Movies") — needed to trigger import scans on held folders
|
||||
HOSTN_LIDARR_DOWNLOAD_DIR="" # host path of the completed-downloads folder Lidarr imports from (e.g. "/mnt/cache/Temp_Storage/SABnzbd/Completed/Music") — blank disables the download orphan cleaner for Lidarr
|
||||
HOSTN_LIDARR_DOWNLOAD_CONTAINER_DIR="" # same folder as Lidarr's container sees it (e.g. "/downloads/Completed/Music") — needed to trigger import scans on held folders
|
||||
|
||||
declare -A HOSTN_RADARR_PATH_MAP=(
|
||||
# ["/movies"]="/mnt/user/Movies"
|
||||
)
|
||||
|
||||
# ━━━ Corruption Scan ━━━
|
||||
# Container that has a working ffprobe binary AND mounts the same shares as the arrs'
|
||||
# roots — check `docker inspect <container>` for its mounts before filling this in.
|
||||
HOSTN_FFPROBE_CONTAINER=""
|
||||
HOSTN_FFPROBE_BIN=""
|
||||
|
||||
declare -A HOSTN_FFPROBE_PATH_MAP=(
|
||||
# ["/mnt/user/Tv_Shows"]="/ext-tv-shows"
|
||||
# ["/mnt/user/Movies"]="/ext-movies"
|
||||
# Prefixes should match what Sonarr/Radarr actually track (HOSTN_SONARR_PATH_MAP /
|
||||
# HOSTN_RADARR_PATH_MAP) — extend coverage here as more shares get mounted into the
|
||||
# ffprobe container, e.g.:
|
||||
# ["/mnt/user/Kids_Tv_Shows"]="/ext-kids-tv"
|
||||
# ["/mnt/user/Kids_Movies"]="/ext-kids-movies"
|
||||
# ["/mnt/user/Anime_Shows"]="/ext-anime-shows"
|
||||
# ["/mnt/user/Anime_Movies"]="/ext-anime-movies"
|
||||
# ["/mnt/user/stand-up_comedy"]="/ext-standup"
|
||||
)
|
||||
|
||||
# ━━━ Arr Recovery Toggles ━━━
|
||||
HOSTN_LIDARR_RECOVERY=false
|
||||
HOSTN_SONARR_RECOVERY=true
|
||||
HOSTN_RADARR_RECOVERY=true
|
||||
|
||||
# ==============================================================================================
|
||||
# ── TRANSCODES ────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
HOSTN_RAMDISK_SIZE="10G"
|
||||
HOSTN_RAMDISK_WARN_GB=8.5
|
||||
HOSTN_RAMDISK_LOW_GB=7
|
||||
HOSTN_TRANSCODE_SSD="/mnt/cache/Temp_Storage/Emby/Transcodes/"
|
||||
|
||||
HOSTN_TRANSCODE_SERVERS=(
|
||||
"${HOSTN_EMBY_CONTAINER}|${HOSTN_EMBY_URL}|${HOSTN_EMBY_API_KEY}|emby"
|
||||
"${HOSTN_JELLYFIN_CONTAINER}|${HOSTN_JELLYFIN_URL}|${HOSTN_JELLYFIN_API_KEY}|jellyfin"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── MONITORS ──────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# ━━━ Certificate Monitor ━━━
|
||||
HOSTN_CERT_MONITOR_DOMAINS=(
|
||||
# "myserver.com"
|
||||
)
|
||||
|
||||
# ━━━ SMART Health ━━━
|
||||
HOSTN_SMART_IGNORE_DRIVES=(
|
||||
"sda" # boot USB — SMART not meaningful on flash drives
|
||||
)
|
||||
|
||||
# ━━━ PCIe AER Quiet ━━━
|
||||
# PCI addresses removed from the bus at array start so dead hardware stops spamming
|
||||
# correctable AER errors into syslog. Full DDDD:BB:DD.F form — find them with:
|
||||
# grep -o "from [0-9a-f:.]*" /var/log/syslog | sort | uniq -c | sort -rn
|
||||
#
|
||||
# Only devices bound to vfio-pci or to no driver at all are eligible. Anything with a
|
||||
# live driver is refused, so a mistyped address cannot pull an HBA or NIC out from
|
||||
# under a running system. Devices claimed by a running VM are refused too.
|
||||
# Gated by PCIE_QUIET_ENABLED in master.conf. Empty list = no-op.
|
||||
HOSTN_PCIE_QUIET_DEVICES=(
|
||||
# "0000:03:00.0"
|
||||
)
|
||||
|
||||
# ━━━ ZFS Report ━━━
|
||||
HOSTN_ZFS_REPORT_IGNORE_POOLS=(
|
||||
# "disk5"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── RESOURCE MANAGER ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
HOSTN_RW_PAUSE_CONTAINERS=(
|
||||
# "Tdarr"
|
||||
# "LidaTube"
|
||||
)
|
||||
|
||||
HOSTN_RW_STOP_CONTAINERS=(
|
||||
# "Tdarr"
|
||||
)
|
||||
|
||||
# ==============================================================================================
|
||||
# ── SYSTEM WATCHDOG ───────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
HOSTN_SYS_WATCHDOG_NIC="" # e.g. eth0 — for network monitoring
|
||||
|
||||
HOSTN_SYS_WATCHDOG_CHECK_DOCKER_DAEMON=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_ROOTFS=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_KERNEL_OOPS=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_FD=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_BOOT=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_OOM=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_RAM=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_LOG=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_ARC=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_CPU_TEMP=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_LOAD=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_ZOMBIES=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_CONTAINERS=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_TMP=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_MDSTAT=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_NETWORK=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_SSHD=true
|
||||
HOSTN_SYS_WATCHDOG_CHECK_RUNAWAY=false
|
||||
|
||||
# ==============================================================================================
|
||||
# ── AUTH STACK ────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Credentials for the Varaverk Auth Stack page (NPM, lldap, Authelia).
|
||||
|
||||
# ━━━ NginxProxyManager ━━━
|
||||
# Admin API runs on 7818 (not 81 — 81 is the partnership WebUI port).
|
||||
HOSTN_NPM_URL="http://localhost:7818"
|
||||
HOSTN_NPM_USER="" # NPM admin email
|
||||
HOSTN_NPM_PASS="" # NPM admin password
|
||||
|
||||
# ━━━ lldap ━━━
|
||||
HOSTN_LLDAP_URL="http://localhost:17170"
|
||||
HOSTN_LLDAP_USER="admin" # lldap admin username
|
||||
HOSTN_LLDAP_PASS="" # lldap admin password
|
||||
|
||||
|
||||
# ==============================================================================================
|
||||
# ── Ollama / AI ───────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Per-host because only some nodes actually have a GPU. A node with an empty OLLAMA_URL is not
|
||||
# an error — it falls through to the resolver and uses another node's Ollama over Tailscale.
|
||||
|
||||
# ━━━ Ollama ━━━
|
||||
HOSTN_OLLAMA_URL="" # e.g. http://localhost:11434 — empty if no local Ollama
|
||||
HOSTN_OLLAMA_GPU_UUID="" # pins Ollama to one card on multi-GPU hosts
|
||||
HOSTN_OLLAMA_MODEL="hf.co/unsloth/Qwen3-14B-GGUF:IQ4_XS" # generation — must fully offload; see README-AI.md
|
||||
|
||||
# ━━━ Web search ━━━
|
||||
# Per-host because one is an address on this network and the other is a credential. Only the
|
||||
# General Chat profile can use these — it is the profile that cannot change anything, which is
|
||||
# why it is the one allowed to look outside. Off until AI_WEB_SEARCH_ENABLED says otherwise.
|
||||
HOSTN_DEGOOG_URL="" # e.g. http://localhost:4444 — self-hosted, no key, /api/search
|
||||
HOSTN_SEARXNG_URL="" # e.g. http://localhost:8888 — needs format: [json] in its settings.yml
|
||||
HOSTN_WEB_SEARCH_API_KEY="" # brave or tavily; unused when the provider is searxng
|
||||
HOSTN_OLLAMA_EMBED_MODEL="nomic-embed-text" # embeddings — the generation model cannot embed
|
||||
|
||||
# ━━━ Authelia ━━━
|
||||
HOSTN_AUTHELIA_CONFIG="/mnt/user/appdata-Fallback/Critical-Data/Authelia/configuration.yml"
|
||||
HOSTN_AUTHELIA_CONTAINER="Authelia"
|
||||
|
||||
# ==============================================================================================
|
||||
# ──────────────────────── End Of HOSTn Variables ──────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
File diff suppressed because it is too large
Load Diff
Executable
+310
@@ -0,0 +1,310 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================== DATA LAYOUT MIGRATION =========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# One-time move of everything Varaverk persists into a single rooted tree under DATA_DIR.
|
||||
#
|
||||
# State_Files/ → data/state/
|
||||
# data/*.db|.count|.tsv|.list → data/db/
|
||||
# data/ai_* → data/ai/
|
||||
# data/*_tracked_cache.json → data/cache/arr/
|
||||
# data/*.log → data/logs/
|
||||
# SCRIPTS_DIR/.cache/vv/d/ → data/cache/conf/
|
||||
#
|
||||
# ==============================================================================================
|
||||
# WHY THIS EXISTS SEPARATELY FROM conf_upgrade
|
||||
# ==============================================================================================
|
||||
#
|
||||
# conf_upgrade adds keys the template has and the installation does not; it never rewrites a
|
||||
# value the operator already has, which is exactly the behaviour you want from it and exactly
|
||||
# why it cannot perform this migration. The paths being moved are existing keys — STATE_DIR,
|
||||
# BANDWIDTH_LOG, AI_INDEX_DB and two dozen more — so their values would keep pointing at the old
|
||||
# layout forever while the new directory variables sat beside them unused.
|
||||
#
|
||||
# So this rewrites those values, then moves the files to match. Both halves, or neither: a conf
|
||||
# pointing at a directory the data is not in is worse than not having started.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Two halves, in order: move the files, then rewrite the conf keys that point at them. Doing it
|
||||
# the other way round would leave every path variable naming a location nothing had reached yet,
|
||||
# and any script that ran in between would create the old layout again underneath the new one.
|
||||
#
|
||||
# Idempotent. A path already under DATA_DIR is left alone, so a re-run after a partial migration
|
||||
# finishes the job rather than moving things twice or failing on what is already done.
|
||||
#
|
||||
# One-time by intent, not by a marker file. There is no "already migrated" flag — the check is
|
||||
# whether each individual path is already where it belongs, which is also what makes an
|
||||
# interrupted run safe to repeat.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Existing keys are rewritten, which is why conf_upgrade cannot do this.
|
||||
# conf_upgrade adds keys the template has and the installation does not, and never rewrites a
|
||||
# value the operator already holds — correct for it, and exactly why it is the wrong tool here.
|
||||
# STATE_DIR, BANDWIDTH_LOG, AI_INDEX_DB and two dozen more are existing keys whose values must
|
||||
# change, or they would go on naming the old layout forever while the new directory variables
|
||||
# sat beside them unused.
|
||||
#
|
||||
# Move, never copy-and-hope.
|
||||
# The data being relocated is the only copy — statistics, histories, the AI index, arr caches.
|
||||
# Everything is moved and the source is gone afterwards, so there is no second location that
|
||||
# might still be written to by something that missed the change.
|
||||
#
|
||||
# The conf rewrite is the last thing, and the riskiest thing.
|
||||
# Until it happens the installation still works from the old layout. That ordering means an
|
||||
# abort partway through leaves a system that runs, rather than one whose paths point at
|
||||
# nothing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Idempotent
|
||||
# Every step tests before acting. A second run reports "already migrated" and changes nothing,
|
||||
# which matters because the natural instinct after a partial failure is to run it again.
|
||||
#
|
||||
# Moves, never copies-and-deletes
|
||||
# mv within one filesystem is atomic per file, so a reader either sees the file at the old
|
||||
# path or the new one — never a half-written copy at both. Nothing is deleted; if a file
|
||||
# cannot be moved it is reported and left exactly where it is.
|
||||
#
|
||||
# Conf is backed up before it is rewritten
|
||||
# master.conf.bak-<stamp>, next to the original, same convention conf_upgrade uses.
|
||||
#
|
||||
# Refuses to run while the orchestrators might be writing
|
||||
# A watchdog that sourced conf before the rewrite and writes state after the move would put a
|
||||
# file back at the old path. The window is seconds and the damage is one stale file, but the
|
||||
# check costs nothing and the failure is silent otherwise.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# This script reads conf to find the old locations and rewrites conf to record the new ones. It
|
||||
# is the one script here whose purpose is to change these values rather than obey them.
|
||||
#
|
||||
# Read to locate what moves
|
||||
# STATE_DIR, BANDWIDTH_LOG, AI_INDEX_DB, AI_MEMORY_FILE, AI_TOKEN_DB, ARR_CLEANUP_STATS,
|
||||
# ARR_SYNC_BLOCKLIST, CORRUPTION_SCAN_STATE_FILE, LIDARR_CACHE_FILE, ZFS_REPORT_LOG and the
|
||||
# rest of the per-script path keys — roughly two dozen in total.
|
||||
#
|
||||
# Written as the new roots
|
||||
# DATA_DIR and the directories beneath it: DB_DIR, STATE_DIR, AI_DATA_DIR,
|
||||
# CACHE_BACKUP_DIR, ARR_CACHE_BACKUP_DIR, CONF_CACHE_BACKUP_DIR, LOG_ARCHIVE_DIR.
|
||||
#
|
||||
# Every rewritten value is expressed as ${DB_DIR}/… rather than an absolute path, so a later
|
||||
# storage-mode migration moves them again by changing one variable.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# migrate_data_layout.sh --dry-run Show what would move. Changes nothing. Do this first.
|
||||
# migrate_data_layout.sh Perform the migration.
|
||||
# migrate_data_layout.sh --force Skip the running-orchestrator check.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
CONF="$ROOT/Configurations/master.conf"
|
||||
|
||||
DRY_RUN=false
|
||||
FORCE=false
|
||||
for a in "$@"; do
|
||||
case "$a" in
|
||||
--dry-run) DRY_RUN=true ;;
|
||||
--force) FORCE=true ;;
|
||||
*) echo "Unknown argument: $a" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
[[ -f "$CONF" ]] || { echo "[FATAL] master.conf not found at $CONF" >&2; exit 1; }
|
||||
|
||||
# Resolve the roots the same way load_config.sh will after this runs.
|
||||
SCRIPTS_DIR="$ROOT"
|
||||
DATA_DIR="$(grep -m1 -E '^\s*DATA_DIR=' "$CONF" | cut -d'"' -f2)"
|
||||
DATA_DIR="${DATA_DIR//\$\{SCRIPTS_DIR\}/$SCRIPTS_DIR}"
|
||||
DATA_DIR="${DATA_DIR:-$ROOT/data}"
|
||||
|
||||
OLD_STATE="$ROOT/State_Files"
|
||||
OLD_CONFCACHE="$ROOT/.cache/vv/d"
|
||||
|
||||
DB_DIR="$DATA_DIR/db"
|
||||
STATE_DIR="$DATA_DIR/state"
|
||||
AI_DATA_DIR="$DATA_DIR/ai"
|
||||
CACHE_BACKUP_DIR="$DATA_DIR/cache"
|
||||
ARR_CACHE_BACKUP_DIR="$CACHE_BACKUP_DIR/arr"
|
||||
CONF_CACHE_BACKUP_DIR="$CACHE_BACKUP_DIR/conf"
|
||||
LOG_ARCHIVE_DIR="$DATA_DIR/logs"
|
||||
|
||||
moved=0; skipped=0; failed=0
|
||||
|
||||
say() { printf ' %s\n' "$*"; }
|
||||
step() { printf '\n━━━ %s ━━━\n' "$*"; }
|
||||
|
||||
# ── Guard: orchestrators mid-run ──────────────────────────────────────────────
|
||||
# Scoped to THIS installation's path, not to the script names. pgrep is system-wide, and a box
|
||||
# running both a production checkout and a development clone will always have one of them busy —
|
||||
# matching on "Orchestrators/" alone made a dev migration abort because prod was mid-cycle, which
|
||||
# is a process that cannot touch this tree's data at all. The full path is what distinguishes
|
||||
# them, so that is what is matched.
|
||||
if [[ "$FORCE" == false && "$DRY_RUN" == false ]]; then
|
||||
running=$(pgrep -fa "$ROOT/(Orchestrators|Watchdogs|Media|Rsync)/" 2>/dev/null \
|
||||
| grep -v "$$" | grep -v migrate_data_layout || true)
|
||||
if [[ -n "$running" ]]; then
|
||||
echo "[ABORT] A job from this installation is running — it may rewrite state mid-move:" >&2
|
||||
echo "$running" >&2
|
||||
echo "Wait for it to finish, or re-run with --force if you are sure." >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Move one path ─────────────────────────────────────────────────────────────
|
||||
move() {
|
||||
local src="$1" dstdir="$2" base
|
||||
base="$(basename "$src")"
|
||||
[[ -e "$src" ]] || return 0
|
||||
if [[ -e "$dstdir/$base" ]]; then
|
||||
say "skip $base — already at ${dstdir#$DATA_DIR/}/"
|
||||
((skipped++)); return 0
|
||||
fi
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
say "would $base → ${dstdir#$DATA_DIR/}/"
|
||||
((moved++)); return 0
|
||||
fi
|
||||
mkdir -p "$dstdir" 2>/dev/null
|
||||
if mv "$src" "$dstdir/$base" 2>/dev/null; then
|
||||
say "moved $base → ${dstdir#$DATA_DIR/}/"
|
||||
((moved++))
|
||||
else
|
||||
say "FAILED $base — left in place"
|
||||
((failed++))
|
||||
fi
|
||||
}
|
||||
|
||||
# ── 1. Rewrite the conf values conf_upgrade cannot ────────────────────────────
|
||||
step "Step 1: master.conf path values"
|
||||
if grep -q 'STATE_DIR="\${DATA_DIR}/state"' "$CONF"; then
|
||||
say "already migrated — no conf changes needed"
|
||||
else
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
cp "$CONF" "${CONF}.bak-$(date +%Y%m%d%H%M%S)"
|
||||
sed -i -E \
|
||||
-e 's|^(\s*STATE_DIR=)".*"|\1"${DATA_DIR}/state"|' \
|
||||
-e 's|^(\s*PERSISTENT_CONF_CACHE=)".*"|\1"${CACHE_BACKUP_DIR}/conf"|' \
|
||||
"$CONF"
|
||||
for v in ARR_SYNC_BLOCKLIST DOCKER_UPDATE_REBUILT_DAILY_FILE DOCKER_UPDATE_REBUILT_WEEKLY_FILE \
|
||||
WATCHDOG_CONTAINER_RESTART_LOG TUNING_MONITOR_LOG LIDARR_TRACKED_COUNT_FILE \
|
||||
LIDARR_RESCAN_DURATION_DB LIDARR_ART_MISS_CACHE LIDARR_DISCOVERY_HISTORY \
|
||||
SONARR_DISCOVERY_HISTORY RADARR_DISCOVERY_HISTORY SONARR_TRACKED_COUNT_FILE \
|
||||
CORRUPTION_SCAN_STATE_FILE CORRUPTION_SCAN_STRIKES_FILE RADARR_TRACKED_COUNT_FILE \
|
||||
TRANSCODE_DAILY_LOG BANDWIDTH_LOG ARR_CLEANUP_STATS ARR_RECOVERY_STATS \
|
||||
ARR_RECOVERY_FAILURE_COUNTS; do
|
||||
sed -i -E "s|^(\s*${v}=\")\\\$\{?DATA_DIR\}?/|\1\${DB_DIR}/|" "$CONF"
|
||||
done
|
||||
for v in AI_INDEX_DB AI_MEMORY_FILE AI_TOKEN_DB; do
|
||||
sed -i -E "s|^(\s*${v}=\")\\\$\{?DATA_DIR\}?/|\1\${AI_DATA_DIR}/|" "$CONF"
|
||||
done
|
||||
sed -i -E 's|^(\s*LIDARR_CACHE_FILE=")\$\{?DATA_DIR\}?/|\1${ARR_CACHE_BACKUP_DIR}/|' "$CONF"
|
||||
sed -i -E 's|^(\s*ZFS_REPORT_LOG=")\$\{?DATA_DIR\}?/|\1${LOG_ARCHIVE_DIR}/|' "$CONF"
|
||||
say "rewritten — backup kept beside it"
|
||||
else
|
||||
say "would rewrite STATE_DIR, PERSISTENT_CONF_CACHE and 25 file paths"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── 2. Build the tree ─────────────────────────────────────────────────────────
|
||||
step "Step 2: directory tree"
|
||||
for d in "$DB_DIR" "$STATE_DIR" "$AI_DATA_DIR" "$ARR_CACHE_BACKUP_DIR" "$LOG_ARCHIVE_DIR"; do
|
||||
if [[ -d "$d" ]]; then say "exists ${d#$DATA_DIR/}"
|
||||
elif [[ "$DRY_RUN" == true ]]; then say "would create ${d#$DATA_DIR/}"
|
||||
else mkdir -p "$d" && say "created ${d#$DATA_DIR/}"
|
||||
fi
|
||||
done
|
||||
# The conf cache carries partner credentials and keeps its restrictive mode.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
mkdir -p "$CONF_CACHE_BACKUP_DIR" && chmod 700 "$CONF_CACHE_BACKUP_DIR"
|
||||
say "created cache/conf (0700)"
|
||||
fi
|
||||
|
||||
# ── 3. State files ────────────────────────────────────────────────────────────
|
||||
step "Step 3: State_Files → data/state"
|
||||
if [[ -d "$OLD_STATE" ]]; then
|
||||
shopt -s nullglob dotglob
|
||||
for f in "$OLD_STATE"/*; do move "$f" "$STATE_DIR"; done
|
||||
shopt -u nullglob dotglob
|
||||
if [[ "$DRY_RUN" == false && -d "$OLD_STATE" ]]; then
|
||||
rmdir "$OLD_STATE" 2>/dev/null && say "removed empty State_Files/" \
|
||||
|| say "State_Files/ not empty — left in place, inspect it"
|
||||
fi
|
||||
else
|
||||
say "no State_Files/ — nothing to do"
|
||||
fi
|
||||
|
||||
# ── 4. Sort the data root ─────────────────────────────────────────────────────
|
||||
step "Step 4: sort data/ into subfolders"
|
||||
classify() {
|
||||
local f="$1" base; base="$(basename "$f")"
|
||||
case "$base" in
|
||||
ai_*) move "$f" "$AI_DATA_DIR" ;;
|
||||
*_tracked_cache.json) move "$f" "$ARR_CACHE_BACKUP_DIR" ;;
|
||||
*.log) move "$f" "$LOG_ARCHIVE_DIR" ;;
|
||||
*.db|*.count|*.tsv|*.list|*.json) move "$f" "$DB_DIR" ;;
|
||||
*) say "leave $base — unclassified, left in data/" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
# Two passes, sidecars first. A SQLite database is three files, and the -wal holds committed
|
||||
# transactions that have not been checkpointed into the .db yet. Move the .db first and any
|
||||
# process that opens it during the gap sees a database with no write-ahead log, creates a fresh
|
||||
# one at the old path, and everything still in the old -wal is lost when it is moved over the
|
||||
# top. Sidecars ahead of their base closes that ordering: the worst case becomes a database
|
||||
# opened without its log still sitting beside it, which SQLite handles.
|
||||
shopt -s nullglob
|
||||
for f in "$DATA_DIR"/*-wal "$DATA_DIR"/*-shm; do
|
||||
[[ -d "$f" ]] && continue
|
||||
classify "$f"
|
||||
done
|
||||
for f in "$DATA_DIR"/*; do
|
||||
# Skip the destinations by name, not everything that happens to be a directory. ai_bugs/ and
|
||||
# ai_chats/ are directories that belong under ai/, and an earlier version of this skipped
|
||||
# every directory outright — which moved neither, silently, while the PHP layer had already
|
||||
# been repointed at the new location. They were empty at the time; that was luck, not design.
|
||||
case "$(basename "$f")" in
|
||||
db|state|ai|cache|logs) continue ;;
|
||||
*-wal|*-shm) continue ;;
|
||||
esac
|
||||
classify "$f"
|
||||
done
|
||||
shopt -u nullglob
|
||||
|
||||
# ── 5. Conf cache backup ──────────────────────────────────────────────────────
|
||||
step "Step 5: conf cache backup"
|
||||
if [[ -d "$OLD_CONFCACHE" ]]; then
|
||||
shopt -s nullglob dotglob
|
||||
for f in "$OLD_CONFCACHE"/*; do move "$f" "$CONF_CACHE_BACKUP_DIR"; done
|
||||
shopt -u nullglob dotglob
|
||||
[[ "$DRY_RUN" == false ]] && rmdir "$OLD_CONFCACHE" "$ROOT/.cache/vv" "$ROOT/.cache" 2>/dev/null
|
||||
else
|
||||
say "no $OLD_CONFCACHE — nothing to do"
|
||||
fi
|
||||
|
||||
# ── Summary ───────────────────────────────────────────────────────────────────
|
||||
printf '\n━━━━━ SUMMARY ━━━━━\n'
|
||||
printf ' %-10s %s\n' "moved:" "$moved"
|
||||
printf ' %-10s %s\n' "skipped:" "$skipped"
|
||||
printf ' %-10s %s\n' "failed:" "$failed"
|
||||
[[ "$DRY_RUN" == true ]] && printf '\n DRY RUN — nothing was changed.\n'
|
||||
[[ "$failed" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
@@ -94,20 +94,54 @@ HOST1_NETWORK_CONNECT_CONTAINERS=(
|
||||
|
||||
```bash
|
||||
# master.conf
|
||||
DAILY_CONTAINER_UPDATES=true # enable/disable daily image pull
|
||||
DAILY_CONTAINER_UPDATES=true # enable/disable the daily image pull
|
||||
# docker_daily_restart.sh still runs regardless
|
||||
# update and restart are independent
|
||||
|
||||
WEEKLY_REMAINING_UPDATES=true # enable/disable weekly remainder pull + prune
|
||||
# to disable: set false or remove from WEEKLY_MAINTENANCE_SCRIPTS
|
||||
WEEKLY_CONTAINER_UPDATES=true # enable/disable the weekly image pull
|
||||
# docker_weekly_restart.sh still runs regardless
|
||||
|
||||
MONTHLY_REMAINING_UPDATES=true # enable/disable the monthly remainder pull
|
||||
# to disable: set false, or remove
|
||||
# "docker_update.sh --remainder" from
|
||||
# MONTHLY_MAINTENANCE_SCRIPTS
|
||||
```
|
||||
|
||||
`docker_update.sh` in normal mode targets `DAILY_RESTART_CONTAINERS` — the same list
|
||||
used by `docker_daily_restart.sh`. No second list to maintain.
|
||||
There is one update script, `docker_update.sh`, with three modes. Each reuses the restart
|
||||
list it pairs with, so there is no second list to maintain:
|
||||
|
||||
`docker_update_remaining.sh` derives its target list automatically:
|
||||
all running containers minus `DAILY_RESTART_CONTAINERS` minus `WEEKLY_RESTART_CONTAINERS`.
|
||||
Everything gets updated at least once per week with no explicit configuration.
|
||||
| Mode | Targets | Runs |
|
||||
|------|---------|------|
|
||||
| *(default)* | `DAILY_RESTART_CONTAINERS` | Daily, before `docker_daily_restart.sh` |
|
||||
| `--weekly` | `WEEKLY_RESTART_CONTAINERS` | Weekly, before `docker_weekly_restart.sh` |
|
||||
| `--remainder` | derived, see below | Monthly, via `MONTHLY_MAINTENANCE_SCRIPTS` |
|
||||
|
||||
Remainder mode needs no configuration at all. It takes every **running** container and
|
||||
subtracts:
|
||||
|
||||
```
|
||||
DAILY_RESTART_CONTAINERS already updated daily
|
||||
WEEKLY_RESTART_CONTAINERS already updated weekly
|
||||
PROFILE_CRITICAL_CONTAINER_NAMES[emby] updated inline by the weekly sync window
|
||||
PROFILE_CRITICAL_CONTAINER_NAMES[critical-data] updated inline by the weekly sync window
|
||||
FALLBACK_<REMOTE>_TIER1..4 owned by the remote's update cycle
|
||||
```
|
||||
|
||||
Stopped containers are never targeted in any mode — pulling for a stopped container adds
|
||||
nothing, and it was most likely stopped deliberately.
|
||||
|
||||
The fallback exclusion is a correctness rule, not an optimisation. This server only runs
|
||||
those containers during a fallback; the remote owns their version. If remainder updated them
|
||||
independently and a handback then occurred, the remote's older image could meet data written
|
||||
by the newer one.
|
||||
|
||||
**Ordering matters.** The update always runs *before* its matching restart so the restart
|
||||
lands on the freshly pulled image. If a container's image actually changed, `docker_update.sh`
|
||||
rebuilds it from its template (a plain `docker restart` reuses the image ID baked in at
|
||||
creation time and would never pick up the new digest) and records it in
|
||||
`DOCKER_UPDATE_REBUILT_*_FILE`. The restart script reads that file and skips those containers
|
||||
rather than restarting them a second time — and discards the file as stale if it is older
|
||||
than `DOCKER_UPDATE_REBUILT_STALE_HOURS`.
|
||||
|
||||
---
|
||||
|
||||
@@ -142,12 +176,22 @@ HOST1_NETWORK_CONNECT_CONTAINERS=()
|
||||
# Shared — applies to both servers
|
||||
|
||||
# ── Container updates ──────────────────────────────────────────────────
|
||||
DAILY_CONTAINER_UPDATES=true
|
||||
WEEKLY_REMAINING_UPDATES=true
|
||||
DAILY_CONTAINER_UPDATES=true # daily pull (DAILY_RESTART_CONTAINERS)
|
||||
WEEKLY_CONTAINER_UPDATES=true # weekly pull (WEEKLY_RESTART_CONTAINERS)
|
||||
MONTHLY_REMAINING_UPDATES=true # monthly pull (everything else)
|
||||
|
||||
# Handoff between update and restart — written by docker_update.sh,
|
||||
# read by the matching restart script so it skips containers already
|
||||
# rebuilt onto a new image this run.
|
||||
DOCKER_UPDATE_REBUILT_DAILY_FILE
|
||||
DOCKER_UPDATE_REBUILT_WEEKLY_FILE
|
||||
DOCKER_UPDATE_REBUILT_STALE_HOURS=12 # older than this = discarded, restart all
|
||||
|
||||
# ── Retry behaviour (shared by restart scripts) ────────────────────────
|
||||
RETRY_COUNT=3 # retry attempts before marking failed
|
||||
SLEEP=5 # seconds between retry attempts
|
||||
CONTAINER_DELAY # seconds between a dependency and its dependents
|
||||
RESTART_VERIFY_WAIT=3 # settle time before verifying a restart stuck
|
||||
|
||||
# Watchdog thresholds (CPU, memory, HTTP, restart loop):
|
||||
# → see Watchdogs/Manual-Watchdogs.md
|
||||
@@ -211,7 +255,24 @@ All scripts support these standard flags:
|
||||
summary block: identity, duration, counts, and a status line. Per-container detail only
|
||||
appears with `--log`. Warnings and errors are always visible regardless of `--log`.
|
||||
|
||||
### `docker_update.sh --remainder`
|
||||
Switches to remainder mode — updates all running containers not in the managed daily/weekly
|
||||
lists. Called by `weekly_sync_maintenance.sh`. Can be run manually to sweep containers
|
||||
that haven't been updated recently.
|
||||
### Script-specific flags
|
||||
|
||||
| Script | Flag | What it does |
|
||||
|--------|------|-------------|
|
||||
| `docker_update.sh` | `--weekly` | Target `WEEKLY_RESTART_CONTAINERS`. Called by `weekly_sync_maintenance.sh` before the weekly restart. |
|
||||
| `docker_update.sh` | `--remainder` | Target every running container not in the daily list, weekly list, emby/critical-data profiles, or fallback tiers. Called by `monthly_maintenance.sh`. Safe to run manually to sweep anything missed. |
|
||||
| `media_cleaner.sh` *(Media/)* | `<anime\|media>` | Required positional profile — there is no default. |
|
||||
|
||||
`docker_update.sh` with no mode flag is normal mode: `DAILY_RESTART_CONTAINERS`, called by
|
||||
`daily_sync_maintenance.sh` before the daily restart.
|
||||
|
||||
### Exit codes
|
||||
|
||||
| Code | Meaning |
|
||||
|------|---------|
|
||||
| `0` | Success, or nothing to do (empty list, disabled toggle, no running containers) |
|
||||
| `1` | One or more containers failed — the summary names them |
|
||||
|
||||
`docker_update.sh` deliberately exits `0` even when pulls fail: a failed pull is not a reason
|
||||
to abort the restart that follows, which simply proceeds on the existing image. The failure
|
||||
is reported in the summary.
|
||||
|
||||
@@ -139,7 +139,7 @@ happens after an update, and you want to know when it does.
|
||||
A container keeps appearing in a broken state. You SSH in and see it's stopped. You
|
||||
don't know if the watchdog tried to restart it and failed, gave up and skip-listed it,
|
||||
is mid-attempt right now, or hasn't noticed yet. You have to manually check the skip
|
||||
list file on `/boot/config/`, check the restart history file, check the state file —
|
||||
list file in `$STATE_DIR`, check the restart history file, check the state file —
|
||||
none of which have obvious formats.
|
||||
|
||||
The fix: `watchdog_skip_list_manager.sh` (in `Tools/`). One command to see exactly what's
|
||||
@@ -153,7 +153,7 @@ its history after you've fixed the problem. No manual file editing required.
|
||||
|
||||
```
|
||||
Docker_Essentials/ ← acts on containers (this folder)
|
||||
unRAID_Essentials/ ← acts on the server itself
|
||||
System_Essentials/ ← acts on the server itself
|
||||
Watchdogs/ ← reactive monitoring + last-resort stability
|
||||
Monitors/ ← observes, measures, reports
|
||||
Rsync/ ← moves data between servers
|
||||
@@ -178,12 +178,28 @@ windows so any downtime from restarts is absorbed by the window that's already h
|
||||
|
||||
---
|
||||
|
||||
### 🔄 Image Currency — `docker_update.sh` + `docker_update_remaining.sh`
|
||||
### 🔄 Image Currency — `docker_update.sh`
|
||||
|
||||
Keeps all container images current without manual intervention. Daily updates for the
|
||||
auth/proxy stack (the containers that restart daily anyway — no extra downtime). Weekly
|
||||
remainder pass for everything else — derives the target list automatically from `docker ps`
|
||||
minus what was already updated, so there is no second list to maintain.
|
||||
One script, three modes — there is no separate remainder script.
|
||||
|
||||
| Mode | Targets | Runs |
|
||||
|------|---------|------|
|
||||
| *(default)* | `DAILY_RESTART_CONTAINERS` | Daily, **before** `docker_daily_restart.sh` |
|
||||
| `--weekly` | `WEEKLY_RESTART_CONTAINERS` | Weekly, **before** `docker_weekly_restart.sh` |
|
||||
| `--remainder` | Everything running that is in neither list | Monthly, via `monthly_maintenance.sh` |
|
||||
|
||||
The update always runs *before* its matching restart, so the restart lands on the freshly
|
||||
pulled image. Reversing that order would restart onto the old image and leave the new one
|
||||
sitting unused until the next window.
|
||||
|
||||
Each mode reuses the restart list it pairs with rather than maintaining its own — add a
|
||||
container to `DAILY_RESTART_CONTAINERS` once and it gets both the restart and the image pull.
|
||||
Remainder mode needs no list at all: it derives its targets from `docker ps` minus the daily
|
||||
list, the weekly list, the emby/critical-data sync-window profiles, and the fallback tiers.
|
||||
|
||||
Fallback containers are deliberately excluded from remainder mode. This server only runs them
|
||||
during a fallback; the remote owns their version. Updating them here would risk the remote's
|
||||
older image meeting data written by a newer one after a handback.
|
||||
|
||||
---
|
||||
|
||||
@@ -207,7 +223,7 @@ or in-progress downloads.
|
||||
## ━━━ RELATIONSHIP TO WATCHDOGS ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
`docker_watchdog.sh` has moved to `Watchdogs/` and is now one of four coordinated
|
||||
single-pass scripts called every minute by `Orchestrators/watchdog_orchestrator.sh`.
|
||||
single-pass scripts called every 15 minutes by `Orchestrators/watchdog_orchestrator.sh`.
|
||||
|
||||
```
|
||||
Watchdogs/resource_watchdog.sh ← reduces pressure before healing attempts
|
||||
@@ -229,8 +245,7 @@ full coordination model between all four watchdogs.
|
||||
|--------|------|-------------|
|
||||
| `docker_daily_restart.sh` | Nightly proactive restart of degradation-prone containers | 1am via `daily_sync_maintenance.sh` |
|
||||
| `docker_weekly_restart.sh` | Weekly restart of less-critical services | 2:30am Sunday via `weekly_sync_maintenance.sh` |
|
||||
| `docker_update.sh` | Container image updates — daily list + weekly remainder mode | Daily before restart; weekly end of window |
|
||||
| `docker_update_remaining.sh` | Image update + prune for all containers not in managed lists | End of weekly maintenance window |
|
||||
| `docker_update.sh` | Container image updates — three modes (default / `--weekly` / `--remainder`) | Daily and weekly before each restart; monthly for the remainder |
|
||||
| `docker_network_connect.sh` | Network existence + container connection enforcement | Every array start |
|
||||
| `docker_container_stop.sh` | Ordered container shutdown — graceful then forced | Called by `array_stopping.sh` |
|
||||
| `downloaders_reset.sh` | Download client hygiene — slskd / SABnzbd / qBittorrent | Every 30min via `critical_sync_maintenance.sh` |
|
||||
@@ -248,7 +263,7 @@ Array starts
|
||||
ensure networks + connections exist
|
||||
silent if correct, notify if creating
|
||||
|
||||
Every minute (Orchestrators/watchdog_orchestrator.sh):
|
||||
Every 15 minutes (Orchestrators/watchdog_orchestrator.sh):
|
||||
└── Watchdogs/docker_watchdog.sh ── see Watchdogs/README-Watchdogs.md
|
||||
|
||||
Daily maintenance window (1am):
|
||||
@@ -258,9 +273,13 @@ Daily maintenance window (1am):
|
||||
|
||||
Weekly maintenance window (2:30am Sunday):
|
||||
weekly_sync_maintenance.sh
|
||||
├── docker_weekly_restart.sh ──── restart less-critical services
|
||||
├── docker_update.sh --remainder ─ update containers not in managed lists
|
||||
└── docker_update_remaining.sh ── prune dangling images
|
||||
├── docker_update.sh --weekly ─── pull latest (WEEKLY_RESTART_CONTAINERS)
|
||||
└── docker_weekly_restart.sh ──── restart onto the fresh image
|
||||
|
||||
Monthly maintenance window:
|
||||
monthly_maintenance.sh
|
||||
├── docker_update.sh --remainder update everything not in the managed lists
|
||||
└── Tools/docker_prune_images.sh --all
|
||||
|
||||
Critical maintenance window (every 30min):
|
||||
critical_sync_maintenance.sh
|
||||
@@ -268,5 +287,39 @@ Critical maintenance window (every 30min):
|
||||
|
||||
Array stopping:
|
||||
array_stopping.sh
|
||||
└── docker_container_stop.sh ──── ordered graceful shutdown
|
||||
└── docker_container_stop.sh ──── ordered graceful shutdown, verified per container
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SAFEGUARDS COMMON TO THIS FOLDER ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
Every script here talks to the Docker daemon, so they share the same protections. Each
|
||||
script's own header documents its full set; these are the ones worth knowing folder-wide.
|
||||
|
||||
**Daemon health is checked, not assumed.** A hung daemon returns an empty container list,
|
||||
which is indistinguishable from "no containers running". Without the check,
|
||||
`docker_container_stop.sh` would report a clean shutdown that never happened, and
|
||||
`docker_update.sh --remainder` would report "nothing to update" while doing nothing.
|
||||
|
||||
**Every docker call is timeout-wrapped.** A wedged daemon cannot stall a maintenance window
|
||||
or hold a lock open. The one deliberate exception is `docker pull` — a large image
|
||||
legitimately outlasts any sane timeout, and killing it mid-layer wastes the transfer.
|
||||
|
||||
**State is respected.** Running containers get restarted; stopped ones stay stopped. A
|
||||
stopped container was almost certainly stopped on purpose, and none of these scripts has the
|
||||
authority to overrule that.
|
||||
|
||||
**Restarts are verified, not assumed.** After each restart the container is re-checked once
|
||||
it has had time to settle. A container that starts and immediately crashes is recorded as a
|
||||
failure and notified — a restart that did not stick is never reported as success.
|
||||
|
||||
**Dependency ordering is shared with the watchdog.** Restarts follow
|
||||
`HOST*_WATCHDOG_DEPENDENCIES`, with `CONTAINER_DELAY` between a dependency and its dependents,
|
||||
so a dependent is never brought up while what it needs is still initialising.
|
||||
|
||||
**Locks prevent overlap.** Long windows can outlast their interval;
|
||||
`downloaders_reset.sh` uses wait-mode because it runs every 30 minutes and the previous pass
|
||||
may still be finishing.
|
||||
|
||||
---
|
||||
|
||||
@@ -68,6 +68,15 @@
|
||||
# Docker Presence Check
|
||||
# Verifies docker binary exists before execution.
|
||||
#
|
||||
# Docker Daemon Check
|
||||
# Verifies the daemon is responsive before enumerating containers. A hung
|
||||
# daemon returns an empty container list, which would otherwise be read as
|
||||
# "nothing to stop" and pass a shutdown that never happened.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() identifies which server is running the script and sets
|
||||
# MY_ID / LOCAL_SERVER_NAME for logging and notifications.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents overlapping runs (e.g. array_stopping firing twice).
|
||||
#
|
||||
@@ -140,17 +149,29 @@ fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be stopped"
|
||||
|
||||
DOCKER_TIMEOUT=30
|
||||
DOCKER_STOP_TIMEOUT=30 # grace period for SIGTERM before docker sends SIGKILL internally
|
||||
|
||||
# A hung daemon makes docker ps return nothing — indistinguishable from "no containers
|
||||
# running", which would silently report a clean shutdown that never happened.
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
|
||||
error "Docker daemon not responding — cannot verify container shutdown"
|
||||
notify "Container stop aborted on $(hostname) ($MY_ID) — Docker daemon not responding" \
|
||||
"Docker Container Stop" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be stopped"
|
||||
|
||||
_RETRY_COUNT="${RETRY_COUNT:-3}"
|
||||
|
||||
log "$ICON_GEAR Config: retries=$_RETRY_COUNT sleep=${SLEEP:-5}s grace=${DOCKER_STOP_TIMEOUT}s cmd-timeout=${DOCKER_TIMEOUT}s"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
mapfile -t RUNNING < <(docker ps --format '{{.Names}}' 2>/dev/null | sort)
|
||||
mapfile -t RUNNING < <(timeout "$DOCKER_TIMEOUT" docker ps --format '{{.Names}}' 2>/dev/null | sort)
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
@@ -164,7 +185,7 @@ fi
|
||||
# ==============================================================================================
|
||||
# ━━━ Stop Containers ━━━
|
||||
# ==============================================================================================
|
||||
mapfile -t RUNNING < <(docker ps --format '{{.Names}}' 2>/dev/null | sort)
|
||||
mapfile -t RUNNING < <(timeout "$DOCKER_TIMEOUT" docker ps --format '{{.Names}}' 2>/dev/null | sort)
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_CONTAINERS Docker Container Stop — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
@@ -175,6 +196,7 @@ if [[ ${#RUNNING[@]} -eq 0 ]]; then
|
||||
fi
|
||||
|
||||
echo "$ICON_CONTAINERS Containers: ${#RUNNING[@]} running"
|
||||
log "$ICON_CONTAINERS Queue: ${RUNNING[*]}"
|
||||
echo ""
|
||||
|
||||
START=$(date +%s)
|
||||
@@ -183,7 +205,9 @@ FAILED=()
|
||||
|
||||
for container in "${RUNNING[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
log "━━━ $ICON_CONTAINERS $container ━━━"
|
||||
c_start=$(date +%s)
|
||||
c_image=$(timeout "$DOCKER_TIMEOUT" docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown")
|
||||
log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would stop $container"
|
||||
@@ -204,7 +228,7 @@ for container in "${RUNNING[@]}"; do
|
||||
-f '{{.State.Running}}' "$container" 2>/dev/null || echo "unknown")
|
||||
|
||||
if [[ "$STATE" != "true" ]]; then
|
||||
log "$ICON_DONE $container stopped ✅"
|
||||
echo "$ICON_DONE $container stopped in $(format_duration $(( $(date +%s) - c_start ))) ✅"
|
||||
STOPPED+=("$container")
|
||||
success=true
|
||||
break
|
||||
@@ -249,7 +273,7 @@ echo "$ICON_CONTAINERS Scope: ${#RUNNING[@]} running → ${#STOPPED[@]} stop
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no containers stopped"
|
||||
elif [[ ${#FAILED[@]} -eq 0 ]]; then
|
||||
log "$ICON_DONE Status: done ✅ — ${#STOPPED[@]} container(s) stopped"
|
||||
echo "$ICON_DONE Status: done ✅ — ${#STOPPED[@]} container(s) stopped"
|
||||
else
|
||||
warn "Status: ${#FAILED[@]} container(s) could not be stopped"
|
||||
fi
|
||||
|
||||
Regular → Executable
+116
-127
@@ -14,6 +14,33 @@
|
||||
# so there is no second list to maintain.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Containers are processed one at a time in dependency-safe order:
|
||||
#
|
||||
# 1. Build restart order
|
||||
# → build_restart_order() sorts DAILY_RESTART_CONTAINERS by WATCHDOG_DEPENDENCIES
|
||||
#
|
||||
# 2. Skip anything docker_update.sh already rebuilt this run
|
||||
# → a rebuild onto a new image already restarted it moments ago
|
||||
#
|
||||
# 3. Inspect container state
|
||||
# missing → skip, not an error
|
||||
# stopped → skip, stopped state is respected
|
||||
# running → restart
|
||||
#
|
||||
# 4. Restart with retry
|
||||
# → retry_docker wraps each attempt in a timeout, up to RETRY_COUNT
|
||||
#
|
||||
# 5. Verify it stayed running
|
||||
# → verify_running() settles for RESTART_VERIFY_WAIT then checks State.Running
|
||||
# → a container that crashes immediately is marked failed and notified
|
||||
#
|
||||
# 6. Prune dangling images
|
||||
# → restarts swap onto new images, leaving the old ones dangling
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -37,6 +64,27 @@
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Docker operations require root privileges.
|
||||
#
|
||||
# Docker Presence Check
|
||||
# Verifies the docker binary exists before execution. Notifies on absence —
|
||||
# a missing binary during the maintenance window is worth knowing about.
|
||||
#
|
||||
# Docker Daemon Check
|
||||
# Verifies the daemon is responsive before any restart work. Every container
|
||||
# would otherwise fail its inspect and be logged as an unknown-status failure,
|
||||
# burying one daemon fault under a list of bogus per-container errors.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() identifies which server is running the script and aliases
|
||||
# HOST*_DAILY_RESTART_CONTAINERS and HOST*_WATCHDOG_DEPENDENCIES to the
|
||||
# correct host's values.
|
||||
#
|
||||
# Empty List Guard
|
||||
# Exits cleanly with a pointer to the relevant conf key if
|
||||
# DAILY_RESTART_CONTAINERS is unconfigured for this host.
|
||||
#
|
||||
# Dependency Ordering
|
||||
# Containers restart in dependency-safe order using HOST*_WATCHDOG_DEPENDENCIES.
|
||||
# CONTAINER_DELAY seconds between dependency restart and dependent restart gives
|
||||
@@ -53,6 +101,11 @@
|
||||
# cannot cause this script to hang indefinitely. Timed-out commands retry
|
||||
# per RETRY_COUNT before marking as failed.
|
||||
#
|
||||
# Stale Rebuild-List Guard
|
||||
# The rebuilt-container list written by docker_update.sh is discarded if older
|
||||
# than DOCKER_UPDATE_REBUILT_STALE_HOURS. A stale file would otherwise suppress
|
||||
# real restarts based on an update run that never happened today.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock() prevents concurrent execution if a previous run is still active.
|
||||
#
|
||||
@@ -82,6 +135,16 @@
|
||||
# CONTAINER_DELAY
|
||||
# Seconds to wait after restarting a dependency before starting its dependents
|
||||
#
|
||||
# RESTART_VERIFY_WAIT
|
||||
# Seconds to wait after docker restart before checking the container is running.
|
||||
# Gives the process time to initialise before verify_running samples the state.
|
||||
# (default: 3)
|
||||
#
|
||||
# DOCKER_UPDATE_REBUILT_DAILY_FILE / DOCKER_UPDATE_REBUILT_STALE_HOURS
|
||||
# List of containers docker_update.sh already rebuilt onto a new image this run —
|
||||
# read here so they're not restarted a second time. Discarded as stale (and every
|
||||
# container restarts normally) if older than DOCKER_UPDATE_REBUILT_STALE_HOURS.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
@@ -124,6 +187,14 @@ fi
|
||||
# detect_hosts() sets MY_ID and aliases HOST*_DAILY_RESTART_CONTAINERS → DAILY_RESTART_CONTAINERS
|
||||
detect_hosts
|
||||
|
||||
# Without this, a hung daemon fails every container's inspect individually and the summary
|
||||
# reports a list of unknown-status failures instead of the one fault that caused them.
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
|
||||
error "Docker daemon not responding — skipping daily restart"
|
||||
notify "Daily restart skipped on $(hostname) — Docker daemon not responding" "Docker Daily Restart" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ${#DAILY_RESTART_CONTAINERS[@]} -eq 0 ]]; then
|
||||
warn "DAILY_RESTART_CONTAINERS is empty for $MY_ID — nothing to restart"
|
||||
warn "Check HOST${MY_ID#HOST}_DAILY_RESTART_CONTAINERS in host*.conf"
|
||||
@@ -152,124 +223,9 @@ fi
|
||||
# ── FUNCTIONS ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Wraps docker commands with a 30 second timeout.
|
||||
# Prevents a hung Docker daemon from causing the script to hang indefinitely.
|
||||
# Usage: docker_cmd docker restart ContainerName
|
||||
DOCKER_TIMEOUT=30
|
||||
docker_cmd() {
|
||||
timeout "$DOCKER_TIMEOUT" "$@"
|
||||
local exit_code=$?
|
||||
if [[ "$exit_code" -eq 124 ]]; then
|
||||
error "Docker command timed out after ${DOCKER_TIMEOUT}s: $*"
|
||||
return 1
|
||||
fi
|
||||
return "$exit_code"
|
||||
}
|
||||
# docker_cmd, verify_running, retry_docker — defined in common.sh
|
||||
|
||||
# Verifies a container is still running after restart.
|
||||
# Gives the container a short settle period before checking.
|
||||
# Returns 0 if running, 1 if crashed or stopped.
|
||||
RESTART_VERIFY_WAIT=5 # seconds to wait before checking state post-restart
|
||||
verify_running() {
|
||||
local container="$1"
|
||||
sleep "$RESTART_VERIFY_WAIT"
|
||||
local state
|
||||
state=$(docker inspect -f '{{.State.Running}}' "$container" 2>/dev/null)
|
||||
if [[ "$state" != "true" ]]; then
|
||||
error "$container failed to stay running after restart — may have crashed"
|
||||
return 1
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# Builds a dependency-safe restart order from DAILY_RESTART_CONTAINERS.
|
||||
# Containers that are dependencies of others restart first.
|
||||
# Returns ordered list in ORDERED_RESTART array.
|
||||
build_restart_order() {
|
||||
ORDERED_RESTART=()
|
||||
local remaining=("${DAILY_RESTART_CONTAINERS[@]}")
|
||||
local placed=()
|
||||
|
||||
# First pass — add dependency containers that appear in our list
|
||||
for container in "${remaining[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
local is_dependency=false
|
||||
# Check if this container is a dependency of any other in our list
|
||||
for dep_string in "${WATCHDOG_DEPENDENCIES[@]}"; do
|
||||
if [[ "$dep_string" == *"$container"* ]]; then
|
||||
is_dependency=true
|
||||
break
|
||||
fi
|
||||
done
|
||||
# Also check associative array format
|
||||
for dependent in "${!WATCHDOG_DEPENDENCIES[@]}"; do
|
||||
if [[ "${WATCHDOG_DEPENDENCIES[$dependent]}" == *"$container"* ]]; then
|
||||
is_dependency=true
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [[ "$is_dependency" == true ]]; then
|
||||
# Check not already placed
|
||||
local already=false
|
||||
for p in "${placed[@]}"; do [[ "$p" == "$container" ]] && already=true && break; done
|
||||
if [[ "$already" == false ]]; then
|
||||
ORDERED_RESTART+=("$container")
|
||||
placed+=("$container")
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
# Second pass — add remaining containers (dependents and independents)
|
||||
for container in "${remaining[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
local already=false
|
||||
for p in "${placed[@]}"; do [[ "$p" == "$container" ]] && already=true && break; done
|
||||
if [[ "$already" == false ]]; then
|
||||
ORDERED_RESTART+=("$container")
|
||||
placed+=("$container")
|
||||
fi
|
||||
done
|
||||
|
||||
log "Restart order: ${ORDERED_RESTART[*]}"
|
||||
}
|
||||
|
||||
# Checks if a container is a dependent of the previously restarted container.
|
||||
# If so, waits CONTAINER_DELAY before restarting to allow dependency to settle.
|
||||
# Usage: check_dependency_delay "$container" "$last_restarted"
|
||||
check_dependency_delay() {
|
||||
local container="$1"
|
||||
local last="$2"
|
||||
[[ -z "$last" ]] && return
|
||||
|
||||
local deps="${WATCHDOG_DEPENDENCIES[$container]:-}"
|
||||
if [[ -n "$deps" ]] && [[ "$deps" == *"$last"* ]]; then
|
||||
echo " Waiting ${CONTAINER_DELAY}s — $container depends on $last..."
|
||||
sleep "$CONTAINER_DELAY"
|
||||
fi
|
||||
}
|
||||
|
||||
# Retries a docker command up to RETRY_COUNT times with SLEEP seconds between attempts.
|
||||
# Uses docker_cmd wrapper for timeout protection on each attempt.
|
||||
# Usage: retry_docker docker restart ContainerName
|
||||
retry_docker() {
|
||||
local attempt=1
|
||||
|
||||
while [[ "$attempt" -le "$RETRY_COUNT" ]]; do
|
||||
log "$ICON_RETRY Attempt $attempt of $RETRY_COUNT: $*"
|
||||
|
||||
if docker_cmd "$@"; then
|
||||
log "Succeeded on attempt $attempt"
|
||||
return 0
|
||||
else
|
||||
warn "Attempt $attempt failed"
|
||||
(( attempt++ ))
|
||||
[[ "$attempt" -le "$RETRY_COUNT" ]] && sleep "$SLEEP"
|
||||
fi
|
||||
done
|
||||
|
||||
error "Command failed after $RETRY_COUNT attempts: $*"
|
||||
return 1
|
||||
}
|
||||
# build_restart_order() / check_dependency_delay() — provided by common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Daily Restart ━━━
|
||||
@@ -277,23 +233,49 @@ retry_docker() {
|
||||
echo ""
|
||||
echo "━━━ $ICON_CONTAINERS Daily Restart — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
echo "$ICON_HOST $MY_ID ($LOCAL_SERVER_NAME) — ${#DAILY_RESTART_CONTAINERS[@]} container(s)"
|
||||
log "$ICON_CONTAINERS Containers: ${DAILY_RESTART_CONTAINERS[*]}"
|
||||
log "$ICON_RETRY Retries: $RETRY_COUNT"
|
||||
log "$ICON_CONTAINERS Containers: ${DAILY_RESTART_CONTAINERS[*]}"
|
||||
log "$ICON_RETRY Retries: $RETRY_COUNT"
|
||||
log "$ICON_GEAR Config: sleep=${SLEEP}s delay=${CONTAINER_DELAY}s verify-wait=${RESTART_VERIFY_WAIT}s cmd-timeout=${DOCKER_TIMEOUT}s"
|
||||
|
||||
START=$(date +%s)
|
||||
FAILED=()
|
||||
RESTARTED=()
|
||||
SKIPPED=()
|
||||
ALREADY_UPDATED=()
|
||||
|
||||
# Build dependency-safe restart order
|
||||
build_restart_order
|
||||
log "$ICON_GEAR Restart order: ${ORDERED_RESTART[*]}"
|
||||
build_restart_order DAILY_RESTART_CONTAINERS
|
||||
|
||||
# ── Load containers docker_update.sh already rebuilt this run ───────────────────────────────────
|
||||
# docker_update.sh's rebuild (stop+recreate onto a new image) already restarted anything whose
|
||||
# image changed today — doing a plain restart on it again here is redundant. A file older than
|
||||
# DOCKER_UPDATE_REBUILT_STALE_HOURS means docker_update.sh either didn't run today or this is way
|
||||
# out of sync with it, so it's discarded rather than trusted, and every container restarts as
|
||||
# normal — same as if the file had never existed.
|
||||
declare -A ALREADY_REBUILT_MAP
|
||||
if [[ -n "${DOCKER_UPDATE_REBUILT_DAILY_FILE:-}" && -f "$DOCKER_UPDATE_REBUILT_DAILY_FILE" ]]; then
|
||||
_rebuilt_age=$(( $(date +%s) - $(stat -c %Y "$DOCKER_UPDATE_REBUILT_DAILY_FILE" 2>/dev/null || echo 0) ))
|
||||
_rebuilt_stale_seconds=$(( ${DOCKER_UPDATE_REBUILT_STALE_HOURS:-12} * 3600 ))
|
||||
if [[ "$_rebuilt_age" -gt "$_rebuilt_stale_seconds" ]]; then
|
||||
warn "Rebuilt-container list is stale ($(( _rebuilt_age / 3600 ))h old) — discarding, restarting all"
|
||||
rm -f "$DOCKER_UPDATE_REBUILT_DAILY_FILE"
|
||||
else
|
||||
while IFS= read -r _c; do
|
||||
[[ -n "$_c" ]] && ALREADY_REBUILT_MAP["$_c"]=1
|
||||
done < "$DOCKER_UPDATE_REBUILT_DAILY_FILE"
|
||||
[[ "${#ALREADY_REBUILT_MAP[@]}" -gt 0 ]] && \
|
||||
log "Already rebuilt today by docker_update.sh, skipping restart: ${!ALREADY_REBUILT_MAP[*]}"
|
||||
fi
|
||||
unset _rebuilt_age _rebuilt_stale_seconds
|
||||
fi
|
||||
|
||||
LAST_RESTARTED=""
|
||||
|
||||
for container in "${ORDERED_RESTART[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
log "━━━ $ICON_CONTAINERS $container ━━━"
|
||||
c_start=$(date +%s)
|
||||
c_image=$(timeout "$DOCKER_TIMEOUT" docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown")
|
||||
log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━"
|
||||
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then
|
||||
warn "$container does not exist — skipping"
|
||||
@@ -304,6 +286,13 @@ for container in "${ORDERED_RESTART[@]}"; do
|
||||
|
||||
case "$STATUS" in
|
||||
true)
|
||||
if [[ -n "${ALREADY_REBUILT_MAP[$container]:-}" ]]; then
|
||||
log "$ICON_RUNNING $container already rebuilt onto new image by docker_update.sh — skipping redundant restart"
|
||||
ALREADY_UPDATED+=("$container")
|
||||
LAST_RESTARTED="$container" # it did restart, just moments ago via the rebuild
|
||||
continue
|
||||
fi
|
||||
|
||||
log "$ICON_RUNNING $container is running — restarting..."
|
||||
|
||||
# Wait if this container depends on the last one restarted
|
||||
@@ -314,9 +303,8 @@ for container in "${ORDERED_RESTART[@]}"; do
|
||||
RESTARTED+=("$container")
|
||||
else
|
||||
if retry_docker docker restart "$container"; then
|
||||
# Verify container stayed running after restart
|
||||
if verify_running "$container"; then
|
||||
log "$ICON_STARTED $container restarted and running ✅"
|
||||
echo "$ICON_STARTED $container restarted and running in $(format_duration $(( $(date +%s) - c_start ))) ✅"
|
||||
RESTARTED+=("$container")
|
||||
LAST_RESTARTED="$container"
|
||||
else
|
||||
@@ -355,7 +343,7 @@ if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would prune dangling images"
|
||||
PRUNED_SUMMARY="(dry run)"
|
||||
else
|
||||
PRUNED_OUTPUT=$(docker image prune -f 2>&1)
|
||||
PRUNED_OUTPUT=$(timeout "$DOCKER_TIMEOUT" docker image prune -f 2>&1)
|
||||
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
|
||||
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
|
||||
fi
|
||||
@@ -366,8 +354,9 @@ fi
|
||||
echo "━━━━━ $ICON_SUMMARY DAILY RESTART SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $((END - START)))"
|
||||
echo "$ICON_CONTAINERS Scope: ${#RESTARTED[@]} restarted, ${#SKIPPED[@]} skipped, ${#FAILED[@]} failed"
|
||||
echo "$ICON_CONTAINERS Scope: ${#RESTARTED[@]} restarted, ${#ALREADY_UPDATED[@]} already updated, ${#SKIPPED[@]} skipped, ${#FAILED[@]} failed"
|
||||
[[ ${#RESTARTED[@]} -gt 0 ]] && log "$ICON_STARTED Restarted: ${RESTARTED[*]}"
|
||||
[[ ${#ALREADY_UPDATED[@]} -gt 0 ]] && log "$ICON_DONE Already updated (skipped): ${ALREADY_UPDATED[*]}"
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "$ICON_NOT_RUNNING Skipped: ${SKIPPED[*]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
echo "$ICON_SYNC Pruned: ${PRUNED_SUMMARY:-none}"
|
||||
@@ -376,7 +365,7 @@ if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ ${#FAILED[@]} -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: ALL DONE ✅"
|
||||
notify "Daily restart complete — ${#RESTARTED[@]} restarted, ${#SKIPPED[@]} skipped on $(hostname)" "Docker Daily Restart" "normal"
|
||||
notify "Daily restart complete — ${#RESTARTED[@]} restarted, ${#ALREADY_UPDATED[@]} already updated, ${#SKIPPED[@]} skipped on $(hostname)" "Docker Daily Restart" "normal"
|
||||
else
|
||||
echo "$ICON_ERROR Status: ${#FAILED[@]} container(s) failed"
|
||||
notify "Daily restart completed with errors on $(hostname) — failed: ${FAILED[*]}" "Docker Daily Restart" "warning"
|
||||
|
||||
@@ -51,6 +51,12 @@
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Docker network operations require root privileges.
|
||||
#
|
||||
# Docker Presence Check
|
||||
# Verifies the docker binary exists before any network operations.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# Prevents concurrent execution via acquire_lock(). Safe to call from
|
||||
# array start hooks or manually without risk of overlap.
|
||||
@@ -71,8 +77,10 @@
|
||||
# Warns and exits cleanly if NETWORK_CONNECT_NETWORKS or
|
||||
# NETWORK_CONNECT_CONTAINERS are unconfigured.
|
||||
#
|
||||
# Command Validation
|
||||
# Validates unRAID notify script before use.
|
||||
# Missing Container Tolerance
|
||||
# A configured container that does not exist yet warns and is skipped rather
|
||||
# than failing the run. This script executes early at array start, before
|
||||
# every container has necessarily been created.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -130,14 +138,12 @@ fi
|
||||
# detect_hosts() sets MY_ID and aliases HOST*_NETWORK_CONNECT_* arrays
|
||||
detect_hosts
|
||||
|
||||
# Validate unRAID notify script — used for network creation alerts
|
||||
validate_unraid_cmd "/usr/local/emhttp/plugins/dynamix/scripts/notify" "" "" "unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
# Docker daemon check — network operations are useless if daemon is hung
|
||||
DOCKER_TIMEOUT=15
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
|
||||
error "Docker daemon not responding — cannot manage networks"
|
||||
notify "docker_network_connect failed on $(hostname) — Docker daemon not responding" "Network Connect" "warning"
|
||||
notify "docker_network_connect failed on $(hostname) — Docker daemon not responding" \
|
||||
"Network Connect" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
@@ -209,7 +215,13 @@ for network in "${NETWORK_CONNECT_NETWORKS[@]}"; do
|
||||
|
||||
# ── Step 1 — ensure network exists ───────────────────────────────────────────────────────
|
||||
if timeout "$DOCKER_TIMEOUT" docker network inspect "$network" &>/dev/null; then
|
||||
log "$network exists ✅"
|
||||
net_subnet=""
|
||||
net_driver=""
|
||||
net_subnet=$(timeout "$DOCKER_TIMEOUT" docker network inspect "$network" \
|
||||
--format '{{range .IPAM.Config}}{{.Subnet}}{{end}}' 2>/dev/null || echo "unknown")
|
||||
net_driver=$(timeout "$DOCKER_TIMEOUT" docker network inspect "$network" \
|
||||
--format '{{.Driver}}' 2>/dev/null || echo "unknown")
|
||||
log "$ICON_DOCKER_NET $network exists ✅ — driver: $net_driver subnet: $net_subnet"
|
||||
else
|
||||
warn "$ICON_DOCKER_NET $network not found — creating (unRAID update may have wiped networks)"
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
@@ -279,14 +291,15 @@ echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $((END - START)))"
|
||||
[[ ${#NETWORKS_CREATED[@]} -gt 0 ]] && warn "$ICON_DOCKER_NET Created: ${NETWORKS_CREATED[*]} (networks were missing)"
|
||||
[[ ${#CONNECTED[@]} -gt 0 ]] && log "Connected: ${CONNECTED[*]}"
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "Already connected: ${#SKIPPED[@]} skipped"
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "Already connected: ${SKIPPED[*]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ ${#FAILED[@]} -gt 0 ]]; then
|
||||
echo "$ICON_ERROR Status: SOME OPERATIONS FAILED"
|
||||
notify "Docker network connect failed on $(hostname) — ${FAILED[*]}" "Network Connect" "warning"
|
||||
notify "Docker network connect failed on $(hostname) — ${FAILED[*]}" \
|
||||
"Network Connect" "warning"
|
||||
elif [[ ${#NETWORKS_CREATED[@]} -gt 0 ]]; then
|
||||
warn "Networks recreated — ${NETWORKS_CREATED[*]} — unRAID update likely wiped them"
|
||||
else
|
||||
|
||||
Regular → Executable
+202
-35
@@ -5,16 +5,18 @@
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Pulls the latest images for configured containers. Two modes: normal (daily)
|
||||
# and remainder (weekly).
|
||||
# Pulls the latest images for configured containers. Three modes:
|
||||
#
|
||||
# Normal mode is called by daily_sync_maintenance.sh before docker_daily_restart.sh.
|
||||
# Normal mode — called by daily_sync_maintenance.sh before docker_daily_restart.sh.
|
||||
# Containers stay running during the pull — no extra downtime beyond what the
|
||||
# nightly restart already causes.
|
||||
#
|
||||
# Remainder mode is called by weekly_sync_maintenance.sh as the final update step.
|
||||
# It catches everything that normal mode and the weekly sync window did not already
|
||||
# update — derived automatically from docker ps, nothing to configure.
|
||||
# Weekly mode — called by weekly_sync_maintenance.sh via WEEKLY_MAINTENANCE_SCRIPTS
|
||||
# before docker_weekly_restart.sh. Pulls WEEKLY_RESTART_CONTAINERS so the
|
||||
# subsequent restart lands on the fresh image.
|
||||
#
|
||||
# Remainder mode — called by monthly_maintenance.sh. Catches everything not
|
||||
# already owned by daily or weekly — derived from docker ps, nothing to configure.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
@@ -25,9 +27,15 @@
|
||||
# Pull → compare old vs new image ID → mark updated or already current.
|
||||
# docker_daily_restart.sh runs after — containers restart onto the fresh image.
|
||||
#
|
||||
# Remainder mode (weekly):
|
||||
# Weekly mode:
|
||||
# Targets WEEKLY_RESTART_CONTAINERS — same list used by docker_weekly_restart.sh.
|
||||
# Pull → compare → rebuild if changed.
|
||||
# docker_weekly_restart.sh runs after — containers restart onto the fresh image.
|
||||
#
|
||||
# Remainder mode (monthly):
|
||||
# Targets all currently running containers NOT in:
|
||||
# DAILY_RESTART_CONTAINERS — already updated daily
|
||||
# WEEKLY_RESTART_CONTAINERS — already updated weekly
|
||||
# emby + critical-data profiles — updated inline by the weekly sync window
|
||||
# FALLBACK_*_TIER* — owned by the remote server's update cycle
|
||||
# Pull → compare → prune dangling images.
|
||||
@@ -43,7 +51,7 @@
|
||||
#
|
||||
# Version Ownership
|
||||
# Fallback containers are excluded from remainder mode. This server only runs
|
||||
# them during a failover. The remote server owns their version — if remainder
|
||||
# them during a fallback. The remote server owns their version — if remainder
|
||||
# updates them independently and a handback occurs, the remote's older image
|
||||
# may not handle data written by the newer version.
|
||||
#
|
||||
@@ -66,13 +74,35 @@
|
||||
# Root Enforcement
|
||||
# Docker operations require root privileges.
|
||||
#
|
||||
# DAILY_CONTAINER_UPDATES Toggle
|
||||
# Normal mode exits cleanly when disabled. docker_daily_restart.sh still runs
|
||||
# regardless — update and restart are independent operations.
|
||||
# Docker Presence Check
|
||||
# Verifies the docker binary exists before execution.
|
||||
#
|
||||
# Docker Daemon Check
|
||||
# Verifies the daemon is responsive before container discovery. Remainder mode
|
||||
# derives its entire target list from docker ps — against a hung daemon that
|
||||
# returns empty and the run silently reports "no containers to update".
|
||||
#
|
||||
# Timeout Protection
|
||||
# Inspect, discovery and image-query commands are wrapped in a timeout so a
|
||||
# hung daemon cannot stall the maintenance window. docker pull is deliberately
|
||||
# NOT wrapped — a large image legitimately takes longer than any sane timeout,
|
||||
# and killing it mid-layer wastes the transfer.
|
||||
#
|
||||
# Empty List Guards
|
||||
# Each mode exits cleanly with a pointer to the relevant conf key when its
|
||||
# container list is unconfigured for this host.
|
||||
#
|
||||
# Rebuild Failure Fallback
|
||||
# A container that fails to rebuild is excluded from the rebuilt-list handoff
|
||||
# file, so the follow-up restart script still gives it a normal restart pass.
|
||||
#
|
||||
# DAILY_CONTAINER_UPDATES / WEEKLY_CONTAINER_UPDATES Toggles
|
||||
# Each mode exits cleanly when disabled. Restart scripts run regardless —
|
||||
# update and restart are independent operations.
|
||||
#
|
||||
# Fallback Exclusion
|
||||
# Remainder mode excludes containers owned by the remote server's update cycle
|
||||
# to prevent version divergence across the failover boundary.
|
||||
# to prevent version divergence across the fallback boundary.
|
||||
#
|
||||
# Running-Only Filter
|
||||
# Stopped containers excluded from remainder mode — intentionally down.
|
||||
@@ -91,6 +121,20 @@
|
||||
# Enable or disable normal mode. docker_daily_restart.sh runs regardless.
|
||||
# (default: true)
|
||||
#
|
||||
# WEEKLY_CONTAINER_UPDATES
|
||||
# Enable or disable weekly mode. docker_weekly_restart.sh runs regardless.
|
||||
# (default: true)
|
||||
#
|
||||
# MONTHLY_REMAINING_UPDATES
|
||||
# Enable or disable remainder mode. (default: true)
|
||||
#
|
||||
# DOCKER_UPDATE_REBUILT_DAILY_FILE / DOCKER_UPDATE_REBUILT_WEEKLY_FILE
|
||||
# Written after each normal/weekly run with the containers actually rebuilt
|
||||
# this pass — read by docker_daily_restart.sh / docker_weekly_restart.sh so
|
||||
# they skip restarting a container a second time right after this script
|
||||
# already rebuilt it onto the new image. Not written in remainder mode
|
||||
# (no follow-up restart script exists for it).
|
||||
#
|
||||
# PROFILE_CRITICAL_CONTAINER_NAMES[emby|critical-data]
|
||||
# Container names for emby and critical-data profiles — excluded from
|
||||
# remainder mode (already updated by the weekly sync window)
|
||||
@@ -101,16 +145,26 @@
|
||||
# Containers updated in normal mode. Aliased by detect_hosts() →
|
||||
# DAILY_RESTART_CONTAINERS
|
||||
#
|
||||
# HOST*_WEEKLY_RESTART_CONTAINERS
|
||||
# Containers updated in weekly mode. Aliased by detect_hosts() →
|
||||
# WEEKLY_RESTART_CONTAINERS
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# docker_update.sh
|
||||
# Normal mode — pull latest images for DAILY_RESTART_CONTAINERS
|
||||
# Called by daily_sync_maintenance.sh before docker_daily_restart.sh
|
||||
#
|
||||
# docker_update.sh --weekly
|
||||
# Weekly mode — pull latest images for WEEKLY_RESTART_CONTAINERS
|
||||
# Called by weekly_sync_maintenance.sh (WEEKLY_MAINTENANCE_SCRIPTS) before docker_weekly_restart.sh
|
||||
#
|
||||
# docker_update.sh --remainder
|
||||
# Remainder mode — pull all running containers not in managed lists,
|
||||
# restart those that received updates, prune dangling images
|
||||
# Called by monthly_maintenance.sh
|
||||
#
|
||||
# docker_update.sh --dry-run
|
||||
# Preview which containers would be pulled without making changes
|
||||
@@ -127,15 +181,16 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# Pre-parse --remainder before parse_args (unknown args pass through to PARSED_ARGS)
|
||||
# Pre-parse mode flags before parse_args (unknown args pass through to PARSED_ARGS)
|
||||
REMAINDER_MODE=false
|
||||
WEEKLY_MODE=false
|
||||
_filtered_args=()
|
||||
for _arg in "$@"; do
|
||||
if [[ "$_arg" == "--remainder" ]]; then
|
||||
REMAINDER_MODE=true
|
||||
else
|
||||
_filtered_args+=("$_arg")
|
||||
fi
|
||||
case "$_arg" in
|
||||
--remainder) REMAINDER_MODE=true ;;
|
||||
--weekly) WEEKLY_MODE=true ;;
|
||||
*) _filtered_args+=("$_arg") ;;
|
||||
esac
|
||||
done
|
||||
unset _arg
|
||||
|
||||
@@ -159,10 +214,23 @@ fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
# Remainder mode builds its whole target list from docker ps — a hung daemon returns
|
||||
# empty and the run would report "no containers to update" instead of failing.
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
|
||||
error "Docker daemon not responding — skipping image updates"
|
||||
notify "Docker update skipped on $(hostname) — Docker daemon not responding" "Docker Update" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Container Discovery ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
if [[ "${MONTHLY_REMAINING_UPDATES:-true}" != "true" ]]; then
|
||||
echo "MONTHLY_REMAINING_UPDATES=false — skipping remainder container updates"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
declare -A _exclude=()
|
||||
|
||||
# Daily containers — updated by docker_update.sh normal mode
|
||||
@@ -170,6 +238,11 @@ if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
[[ -n "$_c" ]] && _exclude["$_c"]=1
|
||||
done
|
||||
|
||||
# Weekly containers — updated by docker_update.sh --weekly
|
||||
for _c in "${WEEKLY_RESTART_CONTAINERS[@]}"; do
|
||||
[[ -n "$_c" ]] && _exclude["$_c"]=1
|
||||
done
|
||||
|
||||
# Weekly sync-window containers (emby + critical-data) — updated inline by weekly_sync_maintenance.sh
|
||||
_weekly_str="${PROFILE_CRITICAL_CONTAINER_NAMES[emby]:-} ${PROFILE_CRITICAL_CONTAINER_NAMES[critical-data]:-}"
|
||||
read -r -a _weekly_arr <<< "$_weekly_str"
|
||||
@@ -179,11 +252,11 @@ if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
unset _weekly_str _weekly_arr
|
||||
|
||||
# Fallback coverage containers — owned by the remote server's update cycle.
|
||||
# This server runs them during failover but should never update them independently.
|
||||
# This server runs them during fallback but should never update them independently.
|
||||
# Updating them here risks version divergence: if remote's writeback after handback
|
||||
# encounters data written by a newer version, it may not handle it correctly.
|
||||
for _tier in 1 2 3 4; do
|
||||
_tier_var="FALLBACK_${MY_ID}_COVERS_${REMOTE_ID}_TIER${_tier}"
|
||||
_tier_var="FALLBACK_${REMOTE_ID}_TIER${_tier}"
|
||||
eval "_tier_arr=(\"\${${_tier_var}[@]:-}\")" 2>/dev/null
|
||||
for _c in "${_tier_arr[@]}"; do
|
||||
[[ -n "$_c" ]] && _exclude["$_c"]=1
|
||||
@@ -191,12 +264,25 @@ if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
done
|
||||
unset _tier _tier_var _tier_arr _c
|
||||
|
||||
mapfile -t _all_running < <(docker ps --format '{{.Names}}' | sort)
|
||||
mapfile -t _all_running < <(timeout "$DOCKER_TIMEOUT" docker ps --format '{{.Names}}' | sort)
|
||||
TARGET_CONTAINERS=()
|
||||
for _c in "${_all_running[@]}"; do
|
||||
[[ -z "${_exclude[$_c]+x}" ]] && TARGET_CONTAINERS+=("$_c")
|
||||
done
|
||||
unset _all_running _exclude _c
|
||||
elif [[ "$WEEKLY_MODE" == true ]]; then
|
||||
if [[ "${WEEKLY_CONTAINER_UPDATES:-true}" != "true" ]]; then
|
||||
echo "WEEKLY_CONTAINER_UPDATES=false — skipping weekly container updates"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ ${#WEEKLY_RESTART_CONTAINERS[@]} -eq 0 ]]; then
|
||||
warn "WEEKLY_RESTART_CONTAINERS is empty for $MY_ID — nothing to update"
|
||||
warn "Check HOST${MY_ID#HOST}_WEEKLY_RESTART_CONTAINERS in host*.conf"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
TARGET_CONTAINERS=("${WEEKLY_RESTART_CONTAINERS[@]}")
|
||||
else
|
||||
if [[ "${DAILY_CONTAINER_UPDATES:-true}" != "true" ]]; then
|
||||
echo "DAILY_CONTAINER_UPDATES=false — skipping container updates"
|
||||
@@ -219,11 +305,16 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Mode: $([[ "$REMAINDER_MODE" == true ]] && echo "remainder" || echo "normal (daily)")"
|
||||
if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
echo "$ICON_CONTAINERS Containers: ${#TARGET_CONTAINERS[@]} running (excluding daily, weekly sync, and fallback)"
|
||||
echo "$ICON_GEAR Mode: remainder (monthly)"
|
||||
echo "$ICON_CONTAINERS Containers: ${#TARGET_CONTAINERS[@]} running (excluding daily, weekly, sync-window, fallback)"
|
||||
for _c in "${TARGET_CONTAINERS[@]}"; do echo " $_c"; done
|
||||
elif [[ "$WEEKLY_MODE" == true ]]; then
|
||||
echo "$ICON_GEAR Mode: weekly"
|
||||
echo "$ICON_CONTAINERS Containers: ${TARGET_CONTAINERS[*]}"
|
||||
echo "$ICON_GEAR Enabled: ${WEEKLY_CONTAINER_UPDATES:-true}"
|
||||
else
|
||||
echo "$ICON_GEAR Mode: normal (daily)"
|
||||
echo "$ICON_CONTAINERS Containers: ${TARGET_CONTAINERS[*]}"
|
||||
echo "$ICON_GEAR Enabled: ${DAILY_CONTAINER_UPDATES:-true}"
|
||||
fi
|
||||
@@ -245,7 +336,11 @@ fi
|
||||
echo ""
|
||||
if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
echo "━━━ $ICON_CONTAINERS Docker Update (remainder) — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
echo "$ICON_CONTAINERS Updating ${#TARGET_CONTAINERS[@]} container(s) (not in daily or weekly sync)"
|
||||
echo "$ICON_CONTAINERS Updating ${#TARGET_CONTAINERS[@]} container(s) (not in daily, weekly, or sync window)"
|
||||
log "$ICON_CONTAINERS Remainder targets: ${TARGET_CONTAINERS[*]}"
|
||||
elif [[ "$WEEKLY_MODE" == true ]]; then
|
||||
echo "━━━ $ICON_CONTAINERS Docker Update (weekly) — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
log "$ICON_CONTAINERS Containers: ${TARGET_CONTAINERS[*]}"
|
||||
else
|
||||
echo "━━━ $ICON_CONTAINERS Docker Update — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
log "$ICON_CONTAINERS Containers: ${TARGET_CONTAINERS[*]}"
|
||||
@@ -256,19 +351,20 @@ START=$(date +%s)
|
||||
UPDATED=()
|
||||
UP_TO_DATE=()
|
||||
FAILED=()
|
||||
OLD_IMAGE_IDS=() # old image IDs to explicitly remove after rebuilds
|
||||
SKIPPED=()
|
||||
|
||||
for container in "${TARGET_CONTAINERS[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
log "━━━ $ICON_CONTAINERS $container ━━━"
|
||||
|
||||
if ! docker inspect "$container" &>/dev/null; then
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then
|
||||
warn "$container — not found, skipping"
|
||||
SKIPPED+=("$container")
|
||||
continue
|
||||
fi
|
||||
|
||||
IMAGE=$(docker inspect --format='{{.Config.Image}}' "$container" 2>/dev/null)
|
||||
IMAGE=$(timeout "$DOCKER_TIMEOUT" docker inspect --format='{{.Config.Image}}' "$container" 2>/dev/null)
|
||||
if [[ -z "$IMAGE" ]]; then
|
||||
warn "$container — could not determine image, skipping"
|
||||
SKIPPED+=("$container")
|
||||
@@ -283,8 +379,11 @@ for container in "${TARGET_CONTAINERS[@]}"; do
|
||||
continue
|
||||
fi
|
||||
|
||||
# Capture image ID before pull to detect whether an update landed
|
||||
OLD_ID=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
# Capture the image ID the container is currently running on, and the
|
||||
# image ID :latest points to before the pull. After pulling, we rebuild if
|
||||
# either a new digest landed OR the container is behind what :latest is now.
|
||||
CONTAINER_IMAGE_ID=$(timeout "$DOCKER_TIMEOUT" docker inspect "$container" --format='{{.Image}}' 2>/dev/null || echo "")
|
||||
OLD_ID=$(timeout "$DOCKER_TIMEOUT" docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
|
||||
log "$ICON_SYNC Pulling $IMAGE..."
|
||||
if [[ "$ENABLE_LOGGING" == "true" ]]; then
|
||||
@@ -294,14 +393,19 @@ for container in "${TARGET_CONTAINERS[@]}"; do
|
||||
docker pull "$IMAGE" >/dev/null 2>&1
|
||||
_pull_rc=$?
|
||||
fi
|
||||
NEW_ID=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
NEW_ID=$(timeout "$DOCKER_TIMEOUT" docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
|
||||
if [[ $_pull_rc -eq 0 ]]; then
|
||||
if [[ -n "$OLD_ID" ]] && [[ "$OLD_ID" != "$NEW_ID" ]]; then
|
||||
log "$ICON_DONE $container — updated ✅"
|
||||
_pull_new=$( [[ -n "$OLD_ID" && "$OLD_ID" != "$NEW_ID" ]] && echo true || echo false)
|
||||
_container_behind=$([[ -n "$CONTAINER_IMAGE_ID" && -n "$NEW_ID" && "$CONTAINER_IMAGE_ID" != "$NEW_ID" ]] && echo true || echo false)
|
||||
|
||||
if [[ "$_pull_new" == true || "$_container_behind" == true ]]; then
|
||||
[[ "$_pull_new" == true ]] && log "$ICON_DONE $container — new image (${OLD_ID:7:12} → ${NEW_ID:7:12})"
|
||||
[[ "$_container_behind" == true && "$_pull_new" == false ]] && log "$ICON_DONE $container — image already pulled, container behind (${CONTAINER_IMAGE_ID:7:12} → ${NEW_ID:7:12})"
|
||||
UPDATED+=("$container")
|
||||
OLD_IMAGE_IDS+=("$CONTAINER_IMAGE_ID")
|
||||
else
|
||||
log "$container — already up to date"
|
||||
log "$container — up to date (${NEW_ID:7:12})"
|
||||
UP_TO_DATE+=("$container")
|
||||
fi
|
||||
else
|
||||
@@ -311,6 +415,64 @@ for container in "${TARGET_CONTAINERS[@]}"; do
|
||||
|
||||
done
|
||||
|
||||
# ── Recreate containers that received a new image ────────────────────────────
|
||||
# docker restart uses the image ID baked in at creation time — it never picks
|
||||
# up the new digest. rebuild_container reads the stored XML template, stops the
|
||||
# old container, recreates it (new image, same config), then prunes the old image.
|
||||
REBUILT=()
|
||||
REBUILD_FAILED=()
|
||||
if [[ ${#UPDATED[@]} -gt 0 ]]; then
|
||||
for container in "${UPDATED[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would rebuild $container from template"
|
||||
REBUILT+=("$container")
|
||||
continue
|
||||
fi
|
||||
log "$ICON_SYNC Rebuilding $container from template on new image..."
|
||||
if platform_rebuild_container "$container"; then
|
||||
echo "$ICON_DONE $container rebuilt ✅"
|
||||
REBUILT+=("$container")
|
||||
else
|
||||
_fallback_msg=$([[ "$WEEKLY_MODE" == true ]] && echo "docker_weekly_restart.sh" || ([[ "$REMAINDER_MODE" == true ]] && echo "no follow-up restart" || echo "docker_daily_restart.sh"))
|
||||
error "Failed to rebuild $container — $_fallback_msg"
|
||||
unset _fallback_msg
|
||||
notify "$container failed to rebuild after image update on $(hostname)" "Docker Update" "warning"
|
||||
REBUILD_FAILED+=("$container")
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
# Record successfully-rebuilt containers so the follow-up restart script (daily/weekly only —
|
||||
# remainder has no follow-up) can skip them instead of restarting an already-fresh container a
|
||||
# second time. Deliberately excludes REBUILD_FAILED — those still need the restart script's
|
||||
# normal pass as a fallback, exactly as the error message above promises. Written even when
|
||||
# REBUILT is empty, so a stale file from a previous run doesn't linger and get misread later.
|
||||
if [[ "$DRY_RUN" == false && "$REMAINDER_MODE" != true ]]; then
|
||||
_rebuilt_file="$DOCKER_UPDATE_REBUILT_DAILY_FILE"
|
||||
[[ "$WEEKLY_MODE" == true ]] && _rebuilt_file="$DOCKER_UPDATE_REBUILT_WEEKLY_FILE"
|
||||
if [[ -n "$_rebuilt_file" ]]; then
|
||||
printf '%s\n' "${REBUILT[@]}" > "$_rebuilt_file" 2>/dev/null
|
||||
fi
|
||||
unset _rebuilt_file
|
||||
fi
|
||||
|
||||
# ── Remove old images ────────────────────────────────────────────────────────
|
||||
# Explicitly rmi by the IDs captured before each pull. Tagged images are never
|
||||
# caught by dangling-only prune, so this is the only reliable cleanup path.
|
||||
# Fall through to dangling prune to catch any leftovers from other update paths.
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would remove ${#OLD_IMAGE_IDS[@]} old image(s) and prune dangling"
|
||||
PRUNED_SUMMARY="(dry run)"
|
||||
else
|
||||
for _old_id in "${OLD_IMAGE_IDS[@]}"; do
|
||||
timeout "$DOCKER_TIMEOUT" docker rmi "$_old_id" >/dev/null 2>&1 || true
|
||||
done
|
||||
PRUNED_OUTPUT=$(timeout "$DOCKER_TIMEOUT" docker image prune -f 2>&1)
|
||||
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
|
||||
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
|
||||
fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -318,6 +480,8 @@ END=$(date +%s)
|
||||
# ==============================================================================================
|
||||
if [[ "$REMAINDER_MODE" == true ]]; then
|
||||
echo "━━━━━ $ICON_SUMMARY DOCKER UPDATE (REMAINDER) SUMMARY ━━━━━"
|
||||
elif [[ "$WEEKLY_MODE" == true ]]; then
|
||||
echo "━━━━━ $ICON_SUMMARY DOCKER UPDATE (WEEKLY) SUMMARY ━━━━━"
|
||||
else
|
||||
echo "━━━━━ $ICON_SUMMARY DOCKER UPDATE SUMMARY ━━━━━"
|
||||
fi
|
||||
@@ -327,9 +491,12 @@ if [[ ${#UPDATED[@]} -gt 0 ]]; then
|
||||
echo "$ICON_DONE Updated: ${#UPDATED[@]}"
|
||||
log " ${UPDATED[*]}"
|
||||
fi
|
||||
[[ ${#UP_TO_DATE[@]} -gt 0 ]] && log "$ICON_RUNNING Up to date: ${#UP_TO_DATE[@]}"
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "$ICON_WARN Skipped: ${#SKIPPED[@]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
[[ ${#REBUILT[@]} -gt 0 ]] && echo "$ICON_SYNC Rebuilt: ${#REBUILT[@]}"
|
||||
[[ ${#REBUILD_FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Rebuild fail:${REBUILD_FAILED[*]}"
|
||||
[[ ${#UP_TO_DATE[@]} -gt 0 ]] && log "$ICON_RUNNING Up to date: ${UP_TO_DATE[*]}"
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "$ICON_WARN Skipped: ${SKIPPED[*]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
echo "$ICON_SYNC Pruned: ${PRUNED_SUMMARY:-none}"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no images pulled"
|
||||
|
||||
@@ -1,351 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ============================= Docker Update — Remaining ======================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Weekly sweep that pulls the latest image for every running container not
|
||||
# already covered by the daily or weekly managed update cycles. Restarts
|
||||
# containers that received a new image, then prunes dangling images.
|
||||
#
|
||||
# Called by weekly_sync_maintenance.sh as the final step in the weekly window.
|
||||
# Derives its target list automatically from docker ps minus the two managed
|
||||
# lists — there is nothing to configure for this script.
|
||||
#
|
||||
# Together with docker_update.sh (normal + remainder modes), every deployed
|
||||
# container receives at least one image pull per week without any per-container
|
||||
# configuration required here.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Lock Acquisition
|
||||
# Prevents concurrent execution via acquire_lock(). Safe to call from
|
||||
# weekly maintenance scripts without risk of overlap.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() identifies which server is running the script and aliases
|
||||
# HOST*_DAILY_RESTART_CONTAINERS and HOST*_WEEKLY_RESTART_CONTAINERS to
|
||||
# the correct host's values for exclusion.
|
||||
#
|
||||
# Root Enforcement
|
||||
# Docker operations require root privileges.
|
||||
#
|
||||
# WEEKLY_REMAINING_UPDATES Toggle
|
||||
# Exits cleanly when disabled via master.conf.
|
||||
#
|
||||
# Running-Only Filter
|
||||
# Stopped containers excluded — intentionally down, pulling adds no value.
|
||||
#
|
||||
# Image ID Comparison
|
||||
# Containers not restarted unless their image actually changed.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# WEEKLY_REMAINING_UPDATES
|
||||
# Enable or disable this script. (default: true)
|
||||
# To disable without the toggle: remove from WEEKLY_MAINTENANCE_SCRIPTS.
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_DAILY_RESTART_CONTAINERS
|
||||
# Excluded from this script — already updated daily. Aliased by
|
||||
# detect_hosts() → DAILY_RESTART_CONTAINERS
|
||||
#
|
||||
# HOST*_WEEKLY_RESTART_CONTAINERS
|
||||
# Excluded from this script — already updated by weekly sync window.
|
||||
# Aliased by detect_hosts() → WEEKLY_RESTART_CONTAINERS
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# docker_update_remaining.sh
|
||||
# Pull all remaining running containers, restart those updated, prune images
|
||||
#
|
||||
# docker_update_remaining.sh --dry-run
|
||||
# Preview which containers would be pulled and restarted
|
||||
#
|
||||
# docker_update_remaining.sh --status
|
||||
# Show exclusion lists and current remaining container count
|
||||
#
|
||||
# docker_update_remaining.sh --log
|
||||
# Verbose per-container pull and restart output
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
if [[ "${WEEKLY_REMAINING_UPDATES:-true}" != "true" ]]; then
|
||||
echo "WEEKLY_REMAINING_UPDATES=false — skipping remaining container updates"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ── Build exclusion set from daily + weekly managed lists ─────────────────────────────────────
|
||||
declare -A EXCLUDED
|
||||
for c in "${DAILY_RESTART_CONTAINERS[@]}" "${WEEKLY_RESTART_CONTAINERS[@]}"; do
|
||||
[[ -n "$c" ]] && EXCLUDED["$c"]=1
|
||||
done
|
||||
|
||||
# ── Get all running containers ────────────────────────────────────────────────────────────────
|
||||
mapfile -t ALL_RUNNING < <(docker ps --format '{{.Names}}' 2>/dev/null | sort)
|
||||
|
||||
# ── Derive remainder: running minus excluded ──────────────────────────────────────────────────
|
||||
REMAINING=()
|
||||
for c in "${ALL_RUNNING[@]}"; do
|
||||
[[ -z "$c" ]] && continue
|
||||
[[ -n "${EXCLUDED[$c]:-}" ]] && continue
|
||||
REMAINING+=("$c")
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Enabled: ${WEEKLY_REMAINING_UPDATES:-true}"
|
||||
echo "$ICON_CONTAINERS All running: ${#ALL_RUNNING[@]}"
|
||||
echo "$ICON_CONTAINERS Excluded: ${!EXCLUDED[*]}"
|
||||
echo "$ICON_CONTAINERS Remaining: ${REMAINING[*]:-none}"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [[ ${#REMAINING[@]} -eq 0 ]]; then
|
||||
echo "No remaining containers to update — all running containers are covered by daily/weekly lists"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no images will be pulled or containers restarted"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FUNCTIONS ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
DOCKER_TIMEOUT=30
|
||||
docker_cmd() {
|
||||
timeout "$DOCKER_TIMEOUT" "$@"
|
||||
local exit_code=$?
|
||||
if [[ "$exit_code" -eq 124 ]]; then
|
||||
error "Docker command timed out after ${DOCKER_TIMEOUT}s: $*"
|
||||
return 1
|
||||
fi
|
||||
return "$exit_code"
|
||||
}
|
||||
|
||||
retry_docker() {
|
||||
local attempt=1
|
||||
while [[ "$attempt" -le "$RETRY_COUNT" ]]; do
|
||||
log "$ICON_RETRY Attempt $attempt of $RETRY_COUNT: $*"
|
||||
if docker_cmd "$@"; then
|
||||
log "Succeeded on attempt $attempt"
|
||||
return 0
|
||||
else
|
||||
warn "Attempt $attempt failed"
|
||||
(( attempt++ ))
|
||||
[[ "$attempt" -le "$RETRY_COUNT" ]] && sleep "$SLEEP"
|
||||
fi
|
||||
done
|
||||
error "Command failed after $RETRY_COUNT attempts: $*"
|
||||
return 1
|
||||
}
|
||||
|
||||
RESTART_VERIFY_WAIT=5
|
||||
verify_running() {
|
||||
local container="$1"
|
||||
sleep "$RESTART_VERIFY_WAIT"
|
||||
local state
|
||||
state=$(docker inspect -f '{{.State.Running}}' "$container" 2>/dev/null)
|
||||
[[ "$state" == "true" ]]
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Pull Updates ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CONTAINERS Docker Update (Remaining) — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
log "$ICON_CONTAINERS Containers: ${REMAINING[*]}"
|
||||
log "$ICON_CONTAINERS Excluded (managed elsewhere): ${!EXCLUDED[*]}"
|
||||
echo ""
|
||||
|
||||
START=$(date +%s)
|
||||
UPDATED=()
|
||||
UP_TO_DATE=()
|
||||
FAILED=()
|
||||
|
||||
for container in "${REMAINING[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
log "━━━ $ICON_CONTAINERS $container ━━━"
|
||||
|
||||
IMAGE=$(docker inspect --format='{{.Config.Image}}' "$container" 2>/dev/null)
|
||||
if [[ -z "$IMAGE" ]]; then
|
||||
warn "$container — could not determine image, skipping"
|
||||
FAILED+=("$container")
|
||||
continue
|
||||
fi
|
||||
|
||||
log "$container — image: $IMAGE"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would pull: $IMAGE"
|
||||
UPDATED+=("$container")
|
||||
continue
|
||||
fi
|
||||
|
||||
OLD_ID=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
|
||||
log "$ICON_SYNC Pulling $IMAGE..."
|
||||
if [[ "$ENABLE_LOGGING" == "true" ]]; then
|
||||
docker pull "$IMAGE" 2>&1 | grep -E "^(Status:|Digest:|Error|error)" | sed 's/^/ /'
|
||||
_pull_rc=${PIPESTATUS[0]}
|
||||
else
|
||||
docker pull "$IMAGE" >/dev/null 2>&1
|
||||
_pull_rc=$?
|
||||
fi
|
||||
NEW_ID=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
|
||||
if [[ $_pull_rc -eq 0 ]]; then
|
||||
if [[ -n "$OLD_ID" ]] && [[ "$OLD_ID" != "$NEW_ID" ]]; then
|
||||
log "$ICON_DONE $container — updated ✅"
|
||||
UPDATED+=("$container")
|
||||
else
|
||||
log "$container — already up to date"
|
||||
UP_TO_DATE+=("$container")
|
||||
fi
|
||||
else
|
||||
warn "$container — pull failed ($IMAGE)"
|
||||
FAILED+=("$container")
|
||||
fi
|
||||
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Restart Updated Containers ━━━
|
||||
# ==============================================================================================
|
||||
RESTARTED=()
|
||||
RESTART_FAILED=()
|
||||
SKIPPED_STOPPED=()
|
||||
|
||||
if [[ ${#UPDATED[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_CONTAINERS Restarting Updated Containers — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
log "$ICON_CONTAINERS Containers with new image: ${UPDATED[*]}"
|
||||
|
||||
for container in "${UPDATED[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
log "━━━ $ICON_CONTAINERS $container ━━━"
|
||||
|
||||
STATUS=$(timeout "$DOCKER_TIMEOUT" docker inspect -f '{{.State.Running}}' "$container" 2>/dev/null)
|
||||
|
||||
if [[ "$STATUS" != "true" ]]; then
|
||||
log "$ICON_NOT_RUNNING $container is stopped — skipping restart (respecting stopped state)"
|
||||
SKIPPED_STOPPED+=("$container")
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would restart $container"
|
||||
RESTARTED+=("$container")
|
||||
continue
|
||||
fi
|
||||
|
||||
log "$ICON_RUNNING $container is running — restarting on new image..."
|
||||
if retry_docker docker restart "$container"; then
|
||||
if verify_running "$container"; then
|
||||
log "$ICON_DONE $container restarted and running ✅"
|
||||
RESTARTED+=("$container")
|
||||
else
|
||||
error "$container restarted but crashed immediately"
|
||||
notify "$container crashed after update-restart on $(hostname)" "Docker Update Remaining" "warning"
|
||||
RESTART_FAILED+=("$container")
|
||||
fi
|
||||
else
|
||||
error "Failed to restart $container after $RETRY_COUNT attempts"
|
||||
notify "$container failed to restart after update on $(hostname)" "Docker Update Remaining" "warning"
|
||||
RESTART_FAILED+=("$container")
|
||||
fi
|
||||
done
|
||||
else
|
||||
log "No containers received a new image — nothing to restart"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Prune Old Images ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Pruning Dangling Images — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would prune dangling images"
|
||||
PRUNED_SUMMARY="(dry run)"
|
||||
else
|
||||
PRUNED_OUTPUT=$(docker image prune -f 2>&1)
|
||||
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
|
||||
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
|
||||
fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY DOCKER UPDATE (REMAINING) SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo "$ICON_CONTAINERS Scope: ${#ALL_RUNNING[@]} running — ${#EXCLUDED[@]} managed = ${#REMAINING[@]} checked"
|
||||
if [[ ${#UPDATED[@]} -gt 0 ]]; then
|
||||
echo "$ICON_DONE New image: ${#UPDATED[@]}"
|
||||
log " ${UPDATED[*]}"
|
||||
fi
|
||||
[[ ${#UP_TO_DATE[@]} -gt 0 ]] && log "$ICON_RUNNING Up to date: ${#UP_TO_DATE[@]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Pull failed: ${FAILED[*]}"
|
||||
if [[ ${#RESTARTED[@]} -gt 0 ]]; then
|
||||
echo "$ICON_DONE Restarted: ${#RESTARTED[@]}"
|
||||
log " ${RESTARTED[*]}"
|
||||
fi
|
||||
[[ ${#SKIPPED_STOPPED[@]} -gt 0 ]] && log "$ICON_WARN Not running: ${#SKIPPED_STOPPED[@]} (skipped restart)"
|
||||
[[ ${#RESTART_FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Restart fail:${RESTART_FAILED[*]}"
|
||||
echo "$ICON_SYNC Pruned: ${PRUNED_SUMMARY:-none}"
|
||||
|
||||
ALL_FAILED=$(( ${#FAILED[@]} + ${#RESTART_FAILED[@]} ))
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ "$ALL_FAILED" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: done ✅ — ${#RESTARTED[@]} restarted, ${#UP_TO_DATE[@]} current"
|
||||
else
|
||||
warn "Status: $ALL_FAILED error(s) — ${#FAILED[@]} pull failure(s), ${#RESTART_FAILED[@]} restart failure(s)"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
[[ "$ALL_FAILED" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
Regular → Executable
+148
-111
@@ -16,9 +16,76 @@
|
||||
# stopped → leave, missing → skip. Container state is always respected.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Identical to docker_daily_restart.sh, against WEEKLY_RESTART_CONTAINERS:
|
||||
#
|
||||
# 1. Build restart order
|
||||
# → build_restart_order() sorts WEEKLY_RESTART_CONTAINERS by WATCHDOG_DEPENDENCIES
|
||||
#
|
||||
# 2. Skip anything docker_update.sh --weekly already rebuilt this run
|
||||
# → a rebuild onto a new image already restarted it moments ago
|
||||
#
|
||||
# 3. Inspect container state
|
||||
# missing → skip, not an error
|
||||
# stopped → skip, stopped state is respected
|
||||
# running → restart
|
||||
#
|
||||
# 4. Restart with retry
|
||||
# → retry_docker wraps each attempt in a timeout, up to RETRY_COUNT
|
||||
#
|
||||
# 5. Verify it stayed running
|
||||
# → verify_running() settles for RESTART_VERIFY_WAIT then checks State.Running
|
||||
# → a container that crashes immediately is marked failed and notified
|
||||
#
|
||||
# 6. Prune dangling images
|
||||
# → restarts swap onto new images, leaving the old ones dangling
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Weekly Cadence
|
||||
# Weekly restarts target services that accumulate state on a slower schedule
|
||||
# than daily targets — less-critical containers that benefit from a periodic
|
||||
# clean start but do not need nightly intervention. Daily restarts handle
|
||||
# high-churn containers; weekly handles the longer-cycle ones.
|
||||
#
|
||||
# State Respect
|
||||
# Running containers are restarted. Stopped containers are left stopped — they
|
||||
# were intentionally halted and this script has no authority to override that
|
||||
# decision. This rule is consistent across the entire ecosystem.
|
||||
#
|
||||
# Dependency-Safe Ordering
|
||||
# Restarts follow the same dependency ordering used by docker_watchdog.sh.
|
||||
# Services that other containers depend on restart first. A dependent is never
|
||||
# restarted while its dependency is still coming up.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Docker operations require root privileges.
|
||||
#
|
||||
# Docker Presence Check
|
||||
# Verifies the docker binary exists before execution. Notifies on absence.
|
||||
#
|
||||
# Docker Daemon Check
|
||||
# Verifies the daemon is responsive before any restart work. Every container
|
||||
# would otherwise fail its inspect and be logged as an unknown-status failure,
|
||||
# burying one daemon fault under a list of bogus per-container errors.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() identifies which server is running the script and aliases
|
||||
# HOST*_WEEKLY_RESTART_CONTAINERS and HOST*_WATCHDOG_DEPENDENCIES to the
|
||||
# correct host's values.
|
||||
#
|
||||
# Empty List Guard
|
||||
# Exits cleanly with a pointer to the relevant conf key if
|
||||
# WEEKLY_RESTART_CONTAINERS is unconfigured for this host.
|
||||
#
|
||||
# Dependency Ordering
|
||||
# Containers restart in dependency-safe order using HOST*_WATCHDOG_DEPENDENCIES.
|
||||
# CONTAINER_DELAY seconds between dependency restart and dependent restart.
|
||||
@@ -31,10 +98,10 @@
|
||||
# All docker commands wrapped in a 30 second timeout. A hung Docker daemon
|
||||
# cannot cause this script to hang indefinitely.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() identifies which server is running the script and aliases
|
||||
# HOST*_WEEKLY_RESTART_CONTAINERS and HOST*_WATCHDOG_DEPENDENCIES to the
|
||||
# correct host's values.
|
||||
# Stale Rebuild-List Guard
|
||||
# The rebuilt-container list written by docker_update.sh --weekly is discarded
|
||||
# if older than DOCKER_UPDATE_REBUILT_STALE_HOURS. A stale file would otherwise
|
||||
# suppress real restarts based on an update run that never happened this week.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock() prevents concurrent execution.
|
||||
@@ -64,6 +131,17 @@
|
||||
# CONTAINER_DELAY
|
||||
# Seconds to wait after restarting a dependency before starting its dependents
|
||||
#
|
||||
# RESTART_VERIFY_WAIT
|
||||
# Seconds verify_running() waits after docker restart before checking the
|
||||
# container is running. Gives the process time to initialise before the
|
||||
# state is sampled. (default: 3)
|
||||
#
|
||||
# DOCKER_UPDATE_REBUILT_WEEKLY_FILE / DOCKER_UPDATE_REBUILT_STALE_HOURS
|
||||
# List of containers docker_update.sh --weekly already rebuilt onto a new image
|
||||
# this run — read here so they're not restarted a second time. Discarded as
|
||||
# stale (and every container restarts normally) if older than
|
||||
# DOCKER_UPDATE_REBUILT_STALE_HOURS.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
@@ -106,6 +184,14 @@ fi
|
||||
# detect_hosts() sets MY_ID and aliases HOST*_WEEKLY_RESTART_CONTAINERS → WEEKLY_RESTART_CONTAINERS
|
||||
detect_hosts
|
||||
|
||||
# Without this, a hung daemon fails every container's inspect individually and the summary
|
||||
# reports a list of unknown-status failures instead of the one fault that caused them.
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker info >/dev/null 2>&1; then
|
||||
error "Docker daemon not responding — skipping weekly restart"
|
||||
notify "Weekly restart skipped on $(hostname) — Docker daemon not responding" "Docker Weekly Restart" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ${#WEEKLY_RESTART_CONTAINERS[@]} -eq 0 ]]; then
|
||||
warn "WEEKLY_RESTART_CONTAINERS is empty for $MY_ID — nothing to restart"
|
||||
warn "Check HOST*_WEEKLY_RESTART_CONTAINERS in host*.conf"
|
||||
@@ -136,108 +222,9 @@ fi
|
||||
# ── FUNCTIONS ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Wraps docker commands with a 30 second timeout.
|
||||
# Prevents a hung Docker daemon from causing the script to hang indefinitely.
|
||||
DOCKER_TIMEOUT=30
|
||||
docker_cmd() {
|
||||
timeout "$DOCKER_TIMEOUT" "$@"
|
||||
local exit_code=$?
|
||||
if [[ "$exit_code" -eq 124 ]]; then
|
||||
error "Docker command timed out after ${DOCKER_TIMEOUT}s: $*"
|
||||
return 1
|
||||
fi
|
||||
return "$exit_code"
|
||||
}
|
||||
# docker_cmd, retry_docker, verify_running — defined in common.sh
|
||||
|
||||
# Retries a docker command up to RETRY_COUNT times with SLEEP seconds between attempts.
|
||||
# Uses docker_cmd wrapper for timeout protection on each attempt.
|
||||
retry_docker() {
|
||||
local attempt=1
|
||||
|
||||
while [[ "$attempt" -le "$RETRY_COUNT" ]]; do
|
||||
log "$ICON_RETRY Attempt $attempt of $RETRY_COUNT: $*"
|
||||
|
||||
if docker_cmd "$@"; then
|
||||
log "Succeeded on attempt $attempt"
|
||||
return 0
|
||||
else
|
||||
warn "Attempt $attempt failed"
|
||||
(( attempt++ ))
|
||||
[[ "$attempt" -le "$RETRY_COUNT" ]] && sleep "$SLEEP"
|
||||
fi
|
||||
done
|
||||
|
||||
error "Command failed after $RETRY_COUNT attempts: $*"
|
||||
return 1
|
||||
}
|
||||
|
||||
# Verifies a container is still running after restart.
|
||||
# Gives the container a short settle period before checking.
|
||||
RESTART_VERIFY_WAIT=5
|
||||
verify_running() {
|
||||
local container="$1"
|
||||
sleep "$RESTART_VERIFY_WAIT"
|
||||
local state
|
||||
state=$(docker inspect -f '{{.State.Running}}' "$container" 2>/dev/null)
|
||||
if [[ "$state" != "true" ]]; then
|
||||
error "$container failed to stay running after restart — may have crashed"
|
||||
return 1
|
||||
fi
|
||||
return 0
|
||||
}
|
||||
|
||||
# Builds a dependency-safe restart order from WEEKLY_RESTART_CONTAINERS.
|
||||
# Containers that are dependencies of others restart first.
|
||||
build_restart_order() {
|
||||
ORDERED_RESTART=()
|
||||
local remaining=("${WEEKLY_RESTART_CONTAINERS[@]}")
|
||||
local placed=()
|
||||
|
||||
# First pass — add dependency containers that appear in our list
|
||||
for container in "${remaining[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
local is_dependency=false
|
||||
for dependent in "${!WATCHDOG_DEPENDENCIES[@]}"; do
|
||||
if [[ "${WATCHDOG_DEPENDENCIES[$dependent]}" == *"$container"* ]]; then
|
||||
is_dependency=true
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [[ "$is_dependency" == true ]]; then
|
||||
local already=false
|
||||
for p in "${placed[@]}"; do [[ "$p" == "$container" ]] && already=true && break; done
|
||||
if [[ "$already" == false ]]; then
|
||||
ORDERED_RESTART+=("$container")
|
||||
placed+=("$container")
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
# Second pass — add remaining containers (dependents and independents)
|
||||
for container in "${remaining[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
local already=false
|
||||
for p in "${placed[@]}"; do [[ "$p" == "$container" ]] && already=true && break; done
|
||||
if [[ "$already" == false ]]; then
|
||||
ORDERED_RESTART+=("$container")
|
||||
placed+=("$container")
|
||||
fi
|
||||
done
|
||||
|
||||
log "Restart order: ${ORDERED_RESTART[*]}"
|
||||
}
|
||||
|
||||
# Waits CONTAINER_DELAY if this container depends on the last restarted one.
|
||||
check_dependency_delay() {
|
||||
local container="$1"
|
||||
local last="$2"
|
||||
[[ -z "$last" ]] && return
|
||||
local deps="${WATCHDOG_DEPENDENCIES[$container]:-}"
|
||||
if [[ -n "$deps" ]] && [[ "$deps" == *"$last"* ]]; then
|
||||
log "Waiting ${CONTAINER_DELAY}s — $container depends on $last..."
|
||||
sleep "$CONTAINER_DELAY"
|
||||
fi
|
||||
}
|
||||
# build_restart_order() / check_dependency_delay() — provided by common.sh
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Weekly Restart ━━━
|
||||
@@ -246,20 +233,46 @@ echo ""
|
||||
echo "━━━ $ICON_CONTAINERS Weekly Restart — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
log "$ICON_CONTAINERS Containers: ${WEEKLY_RESTART_CONTAINERS[*]}"
|
||||
log "$ICON_RETRY Retries: $RETRY_COUNT"
|
||||
log "$ICON_GEAR Config: sleep=${SLEEP}s delay=${CONTAINER_DELAY}s verify-wait=${RESTART_VERIFY_WAIT}s cmd-timeout=${DOCKER_TIMEOUT}s"
|
||||
|
||||
START=$(date +%s)
|
||||
FAILED=()
|
||||
RESTARTED=()
|
||||
SKIPPED=()
|
||||
ALREADY_UPDATED=()
|
||||
|
||||
# Build dependency-safe restart order
|
||||
build_restart_order
|
||||
build_restart_order WEEKLY_RESTART_CONTAINERS
|
||||
|
||||
# ── Load containers docker_update.sh --weekly already rebuilt this run ──────────────────────────
|
||||
# Same reasoning as docker_daily_restart.sh: a container docker_update.sh already rebuilt onto a
|
||||
# new image doesn't need a plain restart right after. A file older than
|
||||
# DOCKER_UPDATE_REBUILT_STALE_HOURS is discarded as untrustworthy rather than trusted, and every
|
||||
# container restarts as normal.
|
||||
declare -A ALREADY_REBUILT_MAP
|
||||
if [[ -n "${DOCKER_UPDATE_REBUILT_WEEKLY_FILE:-}" && -f "$DOCKER_UPDATE_REBUILT_WEEKLY_FILE" ]]; then
|
||||
_rebuilt_age=$(( $(date +%s) - $(stat -c %Y "$DOCKER_UPDATE_REBUILT_WEEKLY_FILE" 2>/dev/null || echo 0) ))
|
||||
_rebuilt_stale_seconds=$(( ${DOCKER_UPDATE_REBUILT_STALE_HOURS:-12} * 3600 ))
|
||||
if [[ "$_rebuilt_age" -gt "$_rebuilt_stale_seconds" ]]; then
|
||||
warn "Rebuilt-container list is stale ($(( _rebuilt_age / 3600 ))h old) — discarding, restarting all"
|
||||
rm -f "$DOCKER_UPDATE_REBUILT_WEEKLY_FILE"
|
||||
else
|
||||
while IFS= read -r _c; do
|
||||
[[ -n "$_c" ]] && ALREADY_REBUILT_MAP["$_c"]=1
|
||||
done < "$DOCKER_UPDATE_REBUILT_WEEKLY_FILE"
|
||||
[[ "${#ALREADY_REBUILT_MAP[@]}" -gt 0 ]] && \
|
||||
log "Already rebuilt today by docker_update.sh, skipping restart: ${!ALREADY_REBUILT_MAP[*]}"
|
||||
fi
|
||||
unset _rebuilt_age _rebuilt_stale_seconds
|
||||
fi
|
||||
|
||||
LAST_RESTARTED=""
|
||||
|
||||
for container in "${ORDERED_RESTART[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
log "━━━ $ICON_CONTAINERS $container ━━━"
|
||||
c_start=$(date +%s)
|
||||
c_image=$(timeout "$DOCKER_TIMEOUT" docker inspect --format '{{.Config.Image}}' "$container" 2>/dev/null || echo "unknown")
|
||||
log "━━━ $ICON_CONTAINERS $container ($c_image) ━━━"
|
||||
|
||||
if ! timeout "$DOCKER_TIMEOUT" docker inspect "$container" &>/dev/null; then
|
||||
warn "$container does not exist — skipping"
|
||||
@@ -270,6 +283,13 @@ for container in "${ORDERED_RESTART[@]}"; do
|
||||
|
||||
case "$STATUS" in
|
||||
true)
|
||||
if [[ -n "${ALREADY_REBUILT_MAP[$container]:-}" ]]; then
|
||||
log "$ICON_RUNNING $container already rebuilt onto new image by docker_update.sh — skipping redundant restart"
|
||||
ALREADY_UPDATED+=("$container")
|
||||
LAST_RESTARTED="$container" # it did restart, just moments ago via the rebuild
|
||||
continue
|
||||
fi
|
||||
|
||||
log "$ICON_RUNNING $container is running — restarting..."
|
||||
|
||||
# Wait if this container depends on the last one restarted
|
||||
@@ -281,7 +301,7 @@ for container in "${ORDERED_RESTART[@]}"; do
|
||||
else
|
||||
if retry_docker docker restart "$container"; then
|
||||
if verify_running "$container"; then
|
||||
log "$ICON_STARTED $container restarted and running ✅"
|
||||
echo "$ICON_STARTED $container restarted and running in $(format_duration $(( $(date +%s) - c_start ))) ✅"
|
||||
RESTARTED+=("$container")
|
||||
LAST_RESTARTED="$container"
|
||||
else
|
||||
@@ -311,6 +331,21 @@ done
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Prune Old Images ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Pruning Dangling Images — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would prune dangling images"
|
||||
PRUNED_SUMMARY="(dry run)"
|
||||
else
|
||||
PRUNED_OUTPUT=$(timeout "$DOCKER_TIMEOUT" docker image prune -f 2>&1)
|
||||
[[ "$ENABLE_LOGGING" == "true" ]] && echo "$PRUNED_OUTPUT" | sed 's/^/ /'
|
||||
PRUNED_SUMMARY=$(echo "$PRUNED_OUTPUT" | grep -E "^Total reclaimed" || echo "nothing reclaimed")
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
@@ -319,16 +354,18 @@ echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $((END - START)))"
|
||||
if [[ ${#RESTARTED[@]} -gt 0 ]]; then
|
||||
echo "$ICON_STARTED Restarted: ${#RESTARTED[@]}"
|
||||
log " ${RESTARTED[*]}"
|
||||
log " Names: ${RESTARTED[*]}"
|
||||
fi
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "$ICON_NOT_RUNNING Skipped: ${#SKIPPED[@]} (were stopped)"
|
||||
[[ ${#ALREADY_UPDATED[@]} -gt 0 ]] && log "$ICON_DONE Already updated (skipped): ${ALREADY_UPDATED[*]}"
|
||||
[[ ${#SKIPPED[@]} -gt 0 ]] && log "$ICON_NOT_RUNNING Skipped: ${SKIPPED[*]} (were stopped)"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
echo "$ICON_SYNC Pruned: ${PRUNED_SUMMARY:-none}"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo "$ICON_WARN Status: DRY RUN — no changes made"
|
||||
elif [[ ${#FAILED[@]} -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: $ICON_SUCCESS ALL DONE"
|
||||
notify "Weekly restart complete — ${#RESTARTED[@]} restarted, ${#SKIPPED[@]} skipped (stopped) on $(hostname)" "Docker Weekly Restart" "normal"
|
||||
notify "Weekly restart complete — ${#RESTARTED[@]} restarted, ${#ALREADY_UPDATED[@]} already updated, ${#SKIPPED[@]} skipped (stopped) on $(hostname)" "Docker Weekly Restart" "normal"
|
||||
else
|
||||
echo "$ICON_ERROR Status: $ICON_ERROR ${#FAILED[@]} container(s) failed"
|
||||
notify "Weekly restart completed with errors on $(hostname) — failed: ${FAILED[*]}" "Docker Weekly Restart" "warning"
|
||||
|
||||
Regular → Executable
+59
-10
@@ -61,6 +61,15 @@
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Failed-import purging deletes directories owned by container users.
|
||||
#
|
||||
# Dependency Check
|
||||
# Verifies curl and jq exist before any API work. jq backs the slskd connection
|
||||
# probe — without it the probe returns false forever, the script burns its full
|
||||
# 60 second reconnect wait, then skips every slskd section as "disconnected".
|
||||
# A missing dependency is reported as itself rather than as a phantom outage.
|
||||
#
|
||||
# Active Transfer Protection
|
||||
# slskd: skips users with InProgress or Queued transfers before any removal.
|
||||
# SABnzbd: age threshold enforced before deletion.
|
||||
@@ -70,6 +79,19 @@
|
||||
# Each section validates its downloader URL before API calls. Missing or
|
||||
# unreachable downloaders skip without affecting other sections.
|
||||
#
|
||||
# No Downloaders Guard
|
||||
# Exits cleanly when none of SLSKD_URL, SABNZBD_URL or QBIT_URL are set for
|
||||
# this host — nothing configured is not an error.
|
||||
#
|
||||
# Timeout Protection
|
||||
# Every curl carries --max-time. An unresponsive downloader cannot stall the
|
||||
# 30 minute maintenance cycle or overlap the next run.
|
||||
#
|
||||
# Deletion Scope Limit
|
||||
# qBittorrent removals pass deleteFiles=false — the torrent record is dropped
|
||||
# but files on disk are left for the arrs to manage. This script never deletes
|
||||
# media.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() identifies which server is running the script and aliases
|
||||
# all HOST*_SLSKD_*, HOST*_SABNZBD_*, and HOST*_QBIT_* vars to the correct
|
||||
@@ -77,7 +99,7 @@
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock "wait" — waits for previous run to finish since this runs every
|
||||
# 15 minutes and prior execution may still be completing.
|
||||
# 30 minutes and prior execution may still be completing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -140,6 +162,17 @@ fi
|
||||
# Lock first — wait mode since this runs every 30min and previous may still be finishing
|
||||
acquire_lock "wait"
|
||||
|
||||
# jq backs the slskd connection probe. Missing, the probe never returns true and slskd
|
||||
# looks permanently disconnected — a 60s wait followed by silently skipped sections.
|
||||
for _dep in curl jq; do
|
||||
if ! command -v "$_dep" &>/dev/null; then
|
||||
error "$_dep not found — required for downloader API calls"
|
||||
notify "Downloaders reset failed on $(hostname) — $_dep not installed" "Downloaders Reset" "warning"
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
unset _dep
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases all HOST*_SLSKD_*, HOST*_SABNZBD_*, HOST*_QBIT_* vars
|
||||
detect_hosts
|
||||
|
||||
@@ -148,6 +181,8 @@ CUTOFF=$(( $(date +%s) - (DOWNLOADER_RETENTION_DAYS * 86400) ))
|
||||
TOTAL_PASS=0
|
||||
TOTAL_FAIL=0
|
||||
|
||||
log "$ICON_GEAR Config: retention=${DOWNLOADER_RETENTION_DAYS}d qbit-age=${QBIT_FAILSAFE_MIN_DAYS}d qbit-ratio=${QBIT_FAILSAFE_MIN_RATIO}"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
@@ -200,7 +235,7 @@ if [[ -n "$SLSKD_URL" ]] && [[ -n "$SLSKD_API_KEY" ]]; then
|
||||
}
|
||||
|
||||
if _slskd_is_connected; then
|
||||
log "slskd connected to Soulseek ✅"
|
||||
echo "$ICON_DONE [OK] slskd connected to Soulseek ✅"
|
||||
SLSKD_CONNECTED=true
|
||||
else
|
||||
warn "slskd disconnected — triggering reconnect"
|
||||
@@ -216,7 +251,7 @@ if [[ -n "$SLSKD_URL" ]] && [[ -n "$SLSKD_API_KEY" ]]; then
|
||||
sleep 10
|
||||
_ELAPSED=$(( _ELAPSED + 10 ))
|
||||
if _slskd_is_connected; then
|
||||
log "slskd reconnected after ${_ELAPSED}s ✅"
|
||||
echo "slskd reconnected after ${_ELAPSED}s ✅"
|
||||
SLSKD_CONNECTED=true
|
||||
break
|
||||
fi
|
||||
@@ -248,7 +283,7 @@ if [[ -n "$SLSKD_URL" ]] && [[ -n "$SLSKD_API_KEY" ]] && [[ "$SLSKD_CONNECTED" =
|
||||
IDS=$(echo "$SEARCHES" | tr '{' '\n' | \
|
||||
grep '"isComplete":true' | grep '"searchText":' | \
|
||||
grep -o '"id":"[^"]*"' | sed 's/"id":"//;s/"//')
|
||||
COUNT=$(echo "$IDS" | grep -c . 2>/dev/null || echo 0)
|
||||
COUNT=$(echo "$IDS" | grep -c . 2>/dev/null || true)
|
||||
COUNT="${COUNT//[^0-9]/}"; COUNT="${COUNT:-0}"
|
||||
|
||||
if [[ "$COUNT" -eq 0 ]]; then
|
||||
@@ -304,6 +339,8 @@ if [[ -n "$SLSKD_URL" ]] && [[ -n "$SLSKD_API_KEY" ]] && [[ "$SLSKD_CONNECTED" =
|
||||
if [[ -z "$USERNAMES" ]]; then
|
||||
success "No transfer records found ✅"
|
||||
else
|
||||
USER_COUNT=$(echo "$USERNAMES" | grep -c . 2>/dev/null || true)
|
||||
log "Found $USER_COUNT user(s) with transfer records"
|
||||
SUCCESS=0; SKIPPED=0; FAIL=0
|
||||
while IFS= read -r USER; do
|
||||
[[ -z "$USER" ]] && continue
|
||||
@@ -379,7 +416,7 @@ if [[ -n "$SLSKD_FAILED_IMPORTS_DIR" ]]; then
|
||||
else
|
||||
OLD_IMPORTS=$(find "$SLSKD_FAILED_IMPORTS_DIR" \
|
||||
-mindepth 1 -maxdepth 1 -mtime +"${DOWNLOADER_RETENTION_DAYS}")
|
||||
IMPORT_COUNT=$(echo "$OLD_IMPORTS" | grep -c . 2>/dev/null || echo 0)
|
||||
IMPORT_COUNT=$(echo "$OLD_IMPORTS" | grep -c . 2>/dev/null || true)
|
||||
IMPORT_COUNT="${IMPORT_COUNT//[^0-9]/}"; IMPORT_COUNT="${IMPORT_COUNT:-0}"
|
||||
|
||||
if [[ "$IMPORT_COUNT" -eq 0 ]]; then
|
||||
@@ -422,6 +459,8 @@ if [[ -n "$SABNZBD_URL" ]] && [[ -n "$SABNZBD_API_KEY" ]]; then
|
||||
if [[ -z "$COMPLETED_IDS" ]]; then
|
||||
success "No completed history found ✅"
|
||||
else
|
||||
HIST_TOTAL=$(echo "$COMPLETED_IDS" | grep -c . 2>/dev/null || true)
|
||||
log "Found $HIST_TOTAL completed history entries"
|
||||
DELETED=0; SKIPPED=0
|
||||
while IFS= read -r NZO_ID; do
|
||||
[[ -z "$NZO_ID" ]] && continue
|
||||
@@ -468,6 +507,8 @@ if [[ -n "$SABNZBD_URL" ]] && [[ -n "$SABNZBD_API_KEY" ]]; then
|
||||
if [[ -z "$FAILED_IDS" ]]; then
|
||||
success "No failed history found ✅"
|
||||
else
|
||||
FAILED_TOTAL=$(echo "$FAILED_IDS" | grep -c . 2>/dev/null || true)
|
||||
log "Found $FAILED_TOTAL failed history entries"
|
||||
DELETED=0; SKIPPED=0
|
||||
while IFS= read -r NZO_ID; do
|
||||
[[ -z "$NZO_ID" ]] && continue
|
||||
@@ -516,6 +557,8 @@ if [[ -n "$SABNZBD_URL" ]] && [[ -n "$SABNZBD_API_KEY" ]]; then
|
||||
if [[ -z "$STALLED_IDS" ]]; then
|
||||
success "No stalled queue items found ✅"
|
||||
else
|
||||
QUEUE_TOTAL=$(echo "$STALLED_IDS" | grep -c . 2>/dev/null || true)
|
||||
log "Found $QUEUE_TOTAL queue item(s) — checking status"
|
||||
DELETED=0; SKIPPED=0
|
||||
while IFS= read -r NZO_ID; do
|
||||
[[ -z "$NZO_ID" ]] && continue
|
||||
@@ -575,6 +618,8 @@ if [[ -n "$QBIT_URL" ]] && [[ -n "$QBIT_USERNAME" ]]; then
|
||||
-H "Cookie: $QBIT_COOKIE" 2>/dev/null)
|
||||
|
||||
NOW=$(date +%s)
|
||||
TORRENT_TOTAL=$(echo "$TORRENTS" | tr '}' '\n' | grep -c '"hash"' 2>/dev/null || true)
|
||||
log "Found $TORRENT_TOTAL torrent(s) — applying age/ratio filter"
|
||||
DELETED=0; SKIPPED=0
|
||||
|
||||
while read -r TORRENT; do
|
||||
@@ -590,11 +635,13 @@ if [[ -n "$QBIT_URL" ]] && [[ -n "$QBIT_USERNAME" ]]; then
|
||||
# Age check — must be old enough
|
||||
[[ "$AGE_DAYS" -lt "$QBIT_FAILSAFE_MIN_DAYS" ]] && ((SKIPPED++)) && continue
|
||||
|
||||
# Ratio check — if configured
|
||||
# Ratio check — if configured. awk handles the fractional comparison;
|
||||
# bash's [[ -lt ]] only does integers and would treat e.g. 1.4 and 1.5
|
||||
# as equal once truncated. Missing/empty ratio defaults to 0 (protected).
|
||||
if [[ "$QBIT_FAILSAFE_MIN_RATIO" != "0" ]]; then
|
||||
RATIO_INT="${RATIO%.*}"
|
||||
MIN_RATIO_INT="${QBIT_FAILSAFE_MIN_RATIO%.*}"
|
||||
[[ "$RATIO_INT" -lt "$MIN_RATIO_INT" ]] && ((SKIPPED++)) && continue
|
||||
if (( $(awk "BEGIN {print (${RATIO:-0} < $QBIT_FAILSAFE_MIN_RATIO) ? 1 : 0}") )); then
|
||||
((SKIPPED++)) && continue
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
@@ -631,8 +678,10 @@ if [[ "$DRY_RUN" == true ]]; then
|
||||
elif [[ "$TOTAL_FAIL" -gt 0 ]]; then
|
||||
echo "$ICON_ERROR Status: $TOTAL_FAIL failure(s) — check logs"
|
||||
notify "Downloaders reset completed with failures on $(hostname)" "Downloaders Reset" "warning"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 1
|
||||
else
|
||||
echo "$ICON_DONE Status: $ICON_SUCCESS DONE"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
+72
-48
@@ -37,11 +37,11 @@ up required before handback sequence starts. **(default: 3)**
|
||||
---
|
||||
|
||||
```
|
||||
FALLBACK_STATE_FILE=/boot/config/fallback_state.db
|
||||
FALLBACK_STATE_FILE="$STATE_DIR/fallback_state.db"
|
||||
```
|
||||
Path to the persistent state file. Lives on `/boot/` intentionally — survives reboots.
|
||||
If the server was in FALLBACK state when it rebooted, it resumes FALLBACK on restart
|
||||
rather than assuming everything is normal.
|
||||
Path to the persistent state file. In `$STATE_DIR` — survives reboots whether storage
|
||||
mode is internal (boot device) or appdata (array). If the server was in FALLBACK state
|
||||
when it rebooted, it resumes FALLBACK on restart rather than assuming everything is normal.
|
||||
|
||||
---
|
||||
|
||||
@@ -105,18 +105,36 @@ internet connectivity. Most hosts leave this empty.
|
||||
---
|
||||
|
||||
```
|
||||
FALLBACK_HOST*_COVERS_HOST*_TIER1=(...)
|
||||
FALLBACK_HOST*_COVERS_HOST*_TIER2=(...)
|
||||
FALLBACK_HOST*_COVERS_HOST*_TIER3=(...)
|
||||
FALLBACK_HOST*_COVERS_HOST*_TIER4=(...)
|
||||
FALLBACK_HOST*_TIER1=(...)
|
||||
FALLBACK_HOST*_TIER2=(...)
|
||||
FALLBACK_HOST*_TIER3=(...)
|
||||
FALLBACK_HOST*_TIER4=(...)
|
||||
```
|
||||
Containers this host starts for the remote host when the remote is down. TIER1 starts
|
||||
immediately. TIER2–4 activate after the corresponding delay thresholds.
|
||||
|
||||
Variable pattern: `FALLBACK_${MY_ID}_COVERS_${REMOTE_ID}_TIER${N}`
|
||||
**Each host declares its OWN services, in its OWN conf.** `FALLBACK_HOST2_TIER1` lives in
|
||||
`host2.conf` and lists HOST2's vital containers — it is not a list HOST1 maintains.
|
||||
|
||||
The DDNS container for the remote's domain must be the first entry in TIER1 — DNS
|
||||
coverage before anything else.
|
||||
At runtime `fallback.sh` reads its **partner's** list:
|
||||
|
||||
```bash
|
||||
get_tier_containers() {
|
||||
local var_name="FALLBACK_${REMOTE_ID}_TIER${tier}" # note: REMOTE_ID, not MY_ID
|
||||
```
|
||||
|
||||
So HOST1, covering HOST2, reads `FALLBACK_HOST2_TIER1` — a variable defined in `host2.conf`
|
||||
and delivered to HOST1 through the partner conf cache (`conf_sync.sh`), because sparse
|
||||
checkout means HOST1 never pulls `host2.conf` from git.
|
||||
|
||||
This is why the naming is what it is. The alternative — each host keeping a copy of its
|
||||
partner's container list — would need editing on both machines every time either one changed
|
||||
a service, and the two copies would silently diverge. Declaring once, on the host that owns
|
||||
the services, means a host is always the authority on what covering it requires.
|
||||
|
||||
TIER1 starts immediately. TIER2–4 activate after their delay thresholds
|
||||
(`HOST*_TIER2_DELAY` and friends, also in that host's own conf).
|
||||
|
||||
The DDNS container for that host's domain must be the **first entry in TIER1** — DNS coverage
|
||||
before anything else.
|
||||
|
||||
---
|
||||
|
||||
@@ -172,7 +190,7 @@ daily_sync_maintenance.sh uses, in the opposite direction. No separate TIER4 lis
|
||||
FALLBACK_ENABLED=true
|
||||
FALLBACK_CHECK_INTERVAL=30
|
||||
FALLBACK_HANDBACK_STRIKES=3
|
||||
FALLBACK_STATE_FILE=/boot/config/fallback_state.db
|
||||
FALLBACK_STATE_FILE="$STATE_DIR/fallback_state.db"
|
||||
FALLBACK_RSYNC_ENABLED=true
|
||||
EXTERNAL_IP=8.8.8.8
|
||||
FALLBACK_TEST_BLOCK_WAIT=60
|
||||
@@ -181,19 +199,19 @@ FALLBACK_TEST_HANDBACK_WAIT=300
|
||||
|
||||
---
|
||||
|
||||
### host2.conf — HOST2 covering HOST1
|
||||
### host1.conf — HOST1's own services (started by HOST2 when HOST1 is down)
|
||||
|
||||
```bash
|
||||
HOST2_DDNS_CONTAINERS=("Gmer4Lfe.us-DDNS")
|
||||
HOST1_DDNS_CONTAINERS=("Gmer4Lfe.com-DDNS")
|
||||
|
||||
FALLBACK_HOST2_STOP_ON_NO_NET=()
|
||||
FALLBACK_HOST1_STOP_ON_NO_NET=()
|
||||
|
||||
# Tier 1 — immediate (vital services + Live TV)
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER1=(
|
||||
FALLBACK_HOST1_TIER1=(
|
||||
"Gmer4Lfe.com-DDNS" # ALWAYS FIRST — DNS coverage before anything else
|
||||
"Emby" # media server — people are watching
|
||||
"NginxProxyManager" # reverse proxy — all external access routes through this
|
||||
"Lldap-Gmer4Lfe" # user directory — already warm, verify and keep
|
||||
"Lldap" # user directory — already warm, verify and keep
|
||||
"Mariadb-Authelia" # auth database — already warm, verify and keep
|
||||
"Redis-Authelia" # auth session cache — already warm, verify and keep
|
||||
"Authelia" # SSO — already warm, serving users already
|
||||
@@ -205,7 +223,7 @@ FALLBACK_HOST2_COVERS_HOST1_TIER1=(
|
||||
)
|
||||
|
||||
# Tier 2 — after 4 hours (shared productivity services)
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER2=(
|
||||
FALLBACK_HOST1_TIER2=(
|
||||
"Postgres-NextCloud" # must start before NextCloud
|
||||
"NextCloud"
|
||||
"PostgreSQL-Immich" # must start before Immich
|
||||
@@ -214,7 +232,7 @@ FALLBACK_HOST2_COVERS_HOST1_TIER2=(
|
||||
)
|
||||
|
||||
# Tier 3 — after 12 hours (secondary services)
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER3=(
|
||||
FALLBACK_HOST1_TIER3=(
|
||||
"Organizrv2-Gmer4Lfe"
|
||||
"AdGuard-Home"
|
||||
"UptimeKuma"
|
||||
@@ -223,7 +241,7 @@ FALLBACK_HOST2_COVERS_HOST1_TIER3=(
|
||||
)
|
||||
|
||||
# Tier 4 — after 24 hours (arrs + downloaders)
|
||||
FALLBACK_HOST2_COVERS_HOST1_TIER4=(
|
||||
FALLBACK_HOST1_TIER4=(
|
||||
"Sonarr-Gmer4Lfe"
|
||||
"Radarr-Gmer4Lfe"
|
||||
"Lidarr-Gmer4Lfe"
|
||||
@@ -259,22 +277,22 @@ FALLBACK_HOST1_WRITEBACK_TIER3=(
|
||||
|
||||
---
|
||||
|
||||
### host1.conf — HOST1 covering HOST2
|
||||
### host2.conf — HOST2's own services (started by HOST1 when HOST2 is down)
|
||||
|
||||
```bash
|
||||
HOST1_DDNS_CONTAINERS=("Gmer4Lfe.com-DDNS")
|
||||
HOST2_DDNS_CONTAINERS=("Gmer4Lfe.us-DDNS")
|
||||
|
||||
FALLBACK_HOST1_STOP_ON_NO_NET=()
|
||||
FALLBACK_HOST2_STOP_ON_NO_NET=()
|
||||
|
||||
# Tier 1 — immediate (HOST2's vital services)
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER1=(
|
||||
FALLBACK_HOST2_TIER1=(
|
||||
"Gmer4Lfe.us-DDNS" # ALWAYS FIRST
|
||||
# HOST2's Tier 1 services — fill per HOST2's stack
|
||||
)
|
||||
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER2=(...)
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER3=(...)
|
||||
FALLBACK_HOST1_COVERS_HOST2_TIER4=(...)
|
||||
FALLBACK_HOST2_TIER2=(...)
|
||||
FALLBACK_HOST2_TIER3=(...)
|
||||
FALLBACK_HOST2_TIER4=(...)
|
||||
|
||||
# Tier delays for HOST2 outage
|
||||
HOST2_TIER2_DELAY=240
|
||||
@@ -293,7 +311,10 @@ FALLBACK_HOST2_WRITEBACK_TIER1=(
|
||||
|
||||
## ━━━ STATE FILE REFERENCE ━━━
|
||||
|
||||
Location: `/boot/config/fallback_state.db` (survives reboots)
|
||||
Location: `$STATE_DIR/fallback_state.db` (survives reboots — boot device or appdata)
|
||||
|
||||
> In a shell where load_config.sh is not sourced, use the full path:
|
||||
> `/boot/config/plugins/varaverk/data/state/fallback_state.db` (internal storage mode)
|
||||
|
||||
```
|
||||
state=NORMAL # NORMAL | FALLBACK | NO_INTERNET | DARK
|
||||
@@ -304,8 +325,8 @@ tier3_started=false # whether Tier 3 containers started
|
||||
tier4_started=false # whether Tier 4 containers started
|
||||
```
|
||||
|
||||
View state: `cat /boot/config/fallback_state.db`
|
||||
Check state: `fallback.sh --status`
|
||||
View state: `fallback.sh --status` (preferred — parsed output)
|
||||
Raw file: `cat "$STATE_DIR/fallback_state.db"` (requires STATE_DIR set, or use full path)
|
||||
|
||||
The file is managed exclusively by fallback.sh. Do not edit it while fallback.sh is
|
||||
running — the next cycle will overwrite your changes. Use the Manual State Reset
|
||||
@@ -352,7 +373,7 @@ ssh root@[HOST2-tailscale-ip] "docker inspect Emby --format '{{.State.Status}}'"
|
||||
|
||||
### 4. Critical Data Mirrored
|
||||
|
||||
These shares must exist on HOST2 with current data from HOST1 before failover is needed:
|
||||
These shares must exist on HOST2 with current data from HOST1 before fallback is needed:
|
||||
|
||||
```
|
||||
/mnt/user/appdata-Fallback/Critical-Data # auth stack — NPM, LLDAP, Authelia, certs
|
||||
@@ -371,7 +392,7 @@ Sync is maintained continuously by `daily_sync_maintenance.sh` critical-data pro
|
||||
### 5. DDNS TTL Set to 1 Minute
|
||||
|
||||
Set in your DDNS provider settings. Higher TTL means users continue hitting the old IP
|
||||
for longer after failover. At 5-minute TTL, users can be hitting a downed server for up
|
||||
for longer after fallback. At 5-minute TTL, users can be hitting a downed server for up
|
||||
to 5 minutes before DNS switches.
|
||||
|
||||
### 6. Both Servers Running fallback.sh
|
||||
@@ -384,7 +405,7 @@ servers must be running it continuously for mutual coverage.
|
||||
pgrep -f "fallback.sh"
|
||||
|
||||
# Check the state file
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
```
|
||||
|
||||
Start via User Scripts plugin on both servers.
|
||||
@@ -398,21 +419,21 @@ current state, outage duration if not NORMAL, Tailscale reachability, and whethe
|
||||
fallback.sh is running.
|
||||
|
||||
**Weekly health digest** (`weekly_health_digest.sh`) — reads the state file. If
|
||||
`DIGEST_SMART_ON_FAILOVER=true` and state is not NORMAL, it sends a notification even
|
||||
`DIGEST_SMART_ON_FALLBACK=true` and state is not NORMAL, it sends a notification even
|
||||
in smart mode — a non-NORMAL state at digest time needs attention.
|
||||
|
||||
**Direct check:**
|
||||
|
||||
```bash
|
||||
fallback.sh --status # full state snapshot
|
||||
cat /boot/config/fallback_state.db # raw state file
|
||||
cat "$STATE_DIR/fallback_state.db" # raw state file
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ PROCEDURES ━━━
|
||||
|
||||
### Running the Failover Test
|
||||
### Running the Fallback Test
|
||||
|
||||
> This starts and stops real containers on both servers. Users will experience a brief
|
||||
> service interruption. Always run `--dry-run` first.
|
||||
@@ -426,7 +447,7 @@ fallback_test.sh
|
||||
|
||||
# Step 3 — check state after test completes
|
||||
fallback.sh --status
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
```
|
||||
|
||||
If the test doesn't complete cleanly, the state file may be left in FALLBACK. The
|
||||
@@ -459,11 +480,14 @@ ping -c 5 [remote-tailscale-ip]
|
||||
**Stop fallback.sh first (via User Scripts Abort), then reset:**
|
||||
|
||||
```bash
|
||||
# Set STATE_DIR (or source load_config.sh to get it from the environment)
|
||||
source /boot/config/plugins/varaverk/load_config.sh
|
||||
|
||||
# View current state
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
|
||||
# Write a clean NORMAL state
|
||||
cat > /boot/config/fallback_state.db << 'EOF'
|
||||
cat > "$STATE_DIR/fallback_state.db" << 'EOF'
|
||||
state=NORMAL
|
||||
fallback_start=0
|
||||
handback_strikes=0
|
||||
@@ -473,7 +497,7 @@ tier4_started=false
|
||||
EOF
|
||||
|
||||
# Verify the write
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
```
|
||||
|
||||
Restart fallback.sh via User Scripts plugin. It will resume from NORMAL on its next cycle.
|
||||
@@ -486,7 +510,7 @@ Restart fallback.sh via User Scripts plugin. It will resume from NORMAL on its n
|
||||
|
||||
1. Create the container on the covering server (stopped), with volume mounts pointing at
|
||||
the mirrored share path (e.g. `/mnt/user/Movies` must exist on the covering server)
|
||||
2. Add the container name to `FALLBACK_HOST*_COVERS_HOST*_TIER*` in host*.conf
|
||||
2. Add the container name to `FALLBACK_<THAT-HOST>_TIER*` in **that host's own** conf
|
||||
in the appropriate tier position (dependency ordering — databases before apps)
|
||||
3. Verify: `fallback.sh --status` shows the container in the expected tier list
|
||||
4. Run `fallback_test.sh --dry-run` to confirm the full configuration is valid
|
||||
@@ -510,7 +534,7 @@ Can this server reach the remote Tailscale IP?
|
||||
|
||||
What does fallback.sh report?
|
||||
→ fallback.sh --status
|
||||
→ cat /boot/config/fallback_state.db
|
||||
→ cat "$STATE_DIR/fallback_state.db"
|
||||
```
|
||||
|
||||
### Handback Not Completing
|
||||
@@ -559,7 +583,7 @@ If it has happened:
|
||||
|
||||
3. Understand the state before resetting
|
||||
fallback.sh --status
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
|
||||
4. Perform Manual State Reset above on the server in a bad state
|
||||
|
||||
@@ -591,9 +615,9 @@ errors are always visible regardless of `--log`.
|
||||
### fallback.sh
|
||||
|
||||
`fallback.sh`
|
||||
Normal start — continuous loop. Start via User Scripts plugin or array_started.sh. **Do NOT
|
||||
stop by killing the process** — state file may be left inconsistent. Stop via User Scripts
|
||||
Abort only.
|
||||
Normal start — continuous loop. Started automatically by `array_started.sh` at array start.
|
||||
**Do NOT stop by killing the process** — state file may be left inconsistent. Stop via
|
||||
`fallback.sh --stop` only.
|
||||
|
||||
`fallback.sh --dry-run`
|
||||
Walk through one full cycle showing what would happen based on current network state. No
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# ━━━━━ FALLBACK ━━━━━
|
||||
|
||||
Mutual automatic failover between two independent unRAID servers. When one goes down the
|
||||
Mutual automatic fallback between two independent unRAID servers. When one goes down the
|
||||
other starts its containers, cuts over DNS, and keeps users online. When it comes back
|
||||
everything hands back in the correct sequence — covering DDNS stops, containers stop, rsync
|
||||
writeback runs, containers start on the primary, primary DDNS starts last — so users hit the
|
||||
@@ -45,7 +45,7 @@ The fix: stop containers before syncing. The outage window is only the rsync dur
|
||||
typically minutes. Clean static source at full bandwidth, predictable state every time.
|
||||
|
||||
**No Way to Validate the System Before Needing It**
|
||||
A failover system that has never been tested is not a failover system — it is a hope.
|
||||
A fallback system that has never been tested is not a fallback system — it is a hope.
|
||||
The fix: `fallback_test.sh` — a controlled simulation using an iptables DROP rule to make
|
||||
the remote appear unreachable, triggering the full sequence without taking anything offline.
|
||||
A safety trap removes the rule on any exit — crash, error, ctrl-c, or clean completion.
|
||||
@@ -115,7 +115,7 @@ Both servers run `fallback.sh` independently as a continuous background process.
|
||||
makes all decisions from two pings every `FALLBACK_CHECK_INTERVAL` seconds:
|
||||
|
||||
```bash
|
||||
ping REMOTE_TAILSCALE_IP # is the other server reachable?
|
||||
ping "$(resolve_tailscale_ip "$REMOTE_SERVER_NAME")" # is the other server reachable?
|
||||
ping EXTERNAL_IP # do I have internet? (default: 8.8.8.8)
|
||||
```
|
||||
|
||||
@@ -220,8 +220,8 @@ determines which server is local and which is remote at runtime, then selects th
|
||||
container arrays and tier delays from config via MY_ID.
|
||||
|
||||
```
|
||||
HOST2 covers HOST1: FALLBACK_HOST2_COVERS_HOST1_TIER* (in host2.conf)
|
||||
HOST1 covers HOST2: FALLBACK_HOST1_COVERS_HOST2_TIER* (in host1.conf)
|
||||
HOST2 covers HOST1: FALLBACK_HOST1_TIER* (in host2.conf)
|
||||
HOST1 covers HOST2: FALLBACK_HOST2_TIER* (in host1.conf)
|
||||
```
|
||||
|
||||
Both servers run identical scripts. MY_ID selects the correct arrays. No hostname
|
||||
|
||||
Executable
+351
@@ -0,0 +1,351 @@
|
||||
#!/bin/bash
|
||||
# ══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
# PURPOSE
|
||||
# Put the containers this host has marked for fallback coverage onto the partner, so that the
|
||||
# partner can actually start them during an outage — and take them off again on request.
|
||||
#
|
||||
# OPERATIONAL MODEL
|
||||
# fallback.sh covers a host by running `docker start <name>` on the partner. It never creates
|
||||
# anything. So a name in FALLBACK_<me>_TIER* is a promise that only holds if the partner already
|
||||
# has that container built. Measured 2026-08-23: all 12 of HOST1's covered containers were absent
|
||||
# from HOST2, meaning every tier would have failed on the first real outage while the UI showed
|
||||
# coverage as configured. This script is what closes that gap.
|
||||
#
|
||||
# Push and remove are separate, deliberate actions, never a side effect of saving the tier list.
|
||||
# Editing coverage is a cheap config write; deploying a dozen containers onto another machine is
|
||||
# not, and the two should not share a button.
|
||||
#
|
||||
# DESIGN PRINCIPLES
|
||||
# Deployed, then verified STOPPED.
|
||||
# A container built here and left running on the partner would be a second live instance of
|
||||
# NextCloud, Gitea or PostgreSQL_Immich against the same data while this host is healthy.
|
||||
# That is the danger_rsync_live_database_appdata failure with worse odds. Every deploy is
|
||||
# followed by a stop and a re-inspect, and a container that will not stay stopped is an
|
||||
# error, not a warning.
|
||||
#
|
||||
# Remove takes the container AND its appdata.
|
||||
# Operator decision 2026-08-23: the button is explicit, so a removal should leave nothing
|
||||
# behind to reason about later. The risk it accepts is narrow and worth naming — if the
|
||||
# partner ever covered for us, ITS appdata is the newer copy and is what a handback rsyncs
|
||||
# home. The NORMAL-state gate below closes the live-failover window; what it cannot see is
|
||||
# a handback that partially failed and then returned to NORMAL, so the UI says so before
|
||||
# asking.
|
||||
#
|
||||
# Two guards on the deletion itself: only paths under /mnt/*/appdata* are ever touched, and
|
||||
# a bind of the appdata ROOT is refused outright — a container mounting /mnt/user/appdata
|
||||
# would otherwise turn one removal into wiping every application on the partner.
|
||||
#
|
||||
# Refuses to run unless fallback state is NORMAL.
|
||||
# Pushing or removing containers mid-outage edits the thing currently keeping services up.
|
||||
#
|
||||
# Coverage names are resolved to templates by <Name>, not by filename.
|
||||
# my-Foo.xml routinely holds a container called something else. Matching on the filename
|
||||
# silently pushes the wrong template, or nothing at all.
|
||||
#
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# Only in NORMAL state. FALLBACK_STATE_FILE is read before anything is pushed or removed, and
|
||||
# any other state refuses the action. A push during a live failover would deploy a second copy
|
||||
# of a container the partner is currently running on our behalf; a remove would delete the one
|
||||
# doing the covering.
|
||||
#
|
||||
# --status is exempt from that gate, because it only reports. Refusing to answer "what is
|
||||
# deployed over there" during a failover would withhold the information precisely when it is
|
||||
# most wanted.
|
||||
#
|
||||
# Every deploy is verified stopped, and a container that will not stay stopped is an error
|
||||
# rather than a warning — see DESIGN PRINCIPLES. A second live instance against the same data
|
||||
# is the failure this whole script exists inside.
|
||||
#
|
||||
# Push and remove are explicit modes with no default. Running the script with no flag does
|
||||
# nothing; neither action can be reached by accident, and neither is a side effect of editing
|
||||
# the tier list.
|
||||
#
|
||||
# --dry-run works in every mode and touches nothing on either host — no container is built,
|
||||
# started, stopped or removed, and no template is written or deleted.
|
||||
#
|
||||
# Remove deletes the container's appdata on the partner as well. That is deliberate and is the
|
||||
# most destructive thing here; the NORMAL-state gate above is what keeps it away from a
|
||||
# partner that is mid-handback.
|
||||
#
|
||||
# CONFIGURATION
|
||||
# master.conf
|
||||
# FALLBACK_<HOST>_TIER1..N the covered container names — what --push deploys and --status
|
||||
# reports on. This script reads that list; it never edits it.
|
||||
#
|
||||
# host*.conf
|
||||
# FALLBACK_STATE_FILE overrides where fallback.sh's state is read from. Defaults to
|
||||
# STATE_DIR/fallback_state.db. A missing file reads as NORMAL,
|
||||
# which is the correct default on a host where fallback has never
|
||||
# run.
|
||||
#
|
||||
# RUNTIME MODES
|
||||
# coverage_deploy.sh --push deploy every covered container onto the partner (stopped)
|
||||
# coverage_deploy.sh --remove stop, remove, and delete the pushed template on the partner
|
||||
# coverage_deploy.sh --status report, per covered container, whether it exists there
|
||||
# any mode supports --dry-run
|
||||
#
|
||||
# DEPENDS ON
|
||||
# Plugin/<platform>/Partnership/containers.sh deploy_container_from_xml(), GPU transform
|
||||
# FALLBACK_<me>_TIER1-4 the coverage list this acts on
|
||||
# ══════════════════════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
source "$SCRIPT_DIR/../Plugin/$PLATFORM/Partnership/containers.sh"
|
||||
|
||||
SSH_TIMEOUT="${SSH_TIMEOUT:-15}"
|
||||
MODE=""
|
||||
DRY_RUN="${DRY_RUN:-false}"
|
||||
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--push) MODE="push" ;;
|
||||
--remove) MODE="remove" ;;
|
||||
--status) MODE="status" ;;
|
||||
--dry-run) DRY_RUN=true ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [[ -z "$MODE" ]]; then
|
||||
error "No mode given — use --push, --remove or --status"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
if [[ -z "$REMOTE_ID" || "$REMOTE_SERVER_NAME" == "unknown" ]]; then
|
||||
error "No partner configured — nothing to push to"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ── Gate: only with fallback idle ─────────────────────────────────────────────────────────────
|
||||
# Read rather than assumed. A missing state file means fallback has never run, which is idle
|
||||
# enough; a file that says anything other than NORMAL means services are in motion right now.
|
||||
FALLBACK_STATE_FILE="${FALLBACK_STATE_FILE:-${STATE_DIR}/fallback_state.db}"
|
||||
_fb_state="NORMAL"
|
||||
if [[ -f "$FALLBACK_STATE_FILE" ]]; then
|
||||
_fb_state=$(grep -m1 '^state=' "$FALLBACK_STATE_FILE" 2>/dev/null | cut -d= -f2)
|
||||
_fb_state="${_fb_state:-NORMAL}"
|
||||
fi
|
||||
if [[ "$_fb_state" != "NORMAL" && "$MODE" != "status" ]]; then
|
||||
error "Fallback state is $_fb_state, not NORMAL — refusing to $MODE"
|
||||
error "Changing what the partner holds while a failover is live edits the thing keeping services up."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ── The coverage list ─────────────────────────────────────────────────────────────────────────
|
||||
COVERED=()
|
||||
for _t in 1 2 3 4; do
|
||||
_var="FALLBACK_${MY_ID}_TIER${_t}[@]"
|
||||
for _c in "${!_var}"; do
|
||||
[[ -n "$_c" ]] && COVERED+=("$_c")
|
||||
done
|
||||
done
|
||||
|
||||
if [[ ${#COVERED[@]} -eq 0 ]]; then
|
||||
warn "No containers are covered in FALLBACK_${MY_ID}_TIER1-4 — nothing to do"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
log "$ICON_FALLBACK Coverage: ${#COVERED[@]} container(s) for $REMOTE_SERVER_NAME to start during an outage"
|
||||
|
||||
resolve_remote_ip
|
||||
MIRROR="$REMOTE_SERVER_NAME"
|
||||
MIRROR_IP="$REMOTE_SERVER"
|
||||
_key_var="${MY_ID}_SSH_KEY"
|
||||
MIRROR_SSH_KEY="${!_key_var}"
|
||||
|
||||
if [[ ! -f "$MIRROR_SSH_KEY" ]]; then
|
||||
error "SSH key $MIRROR_SSH_KEY not found — cannot reach $MIRROR"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ── name -> template ──────────────────────────────────────────────────────────────────────────
|
||||
# Matched on the <Name> element. Filenames lie often enough that trusting them would push the
|
||||
# wrong container without saying so.
|
||||
xml_for_container() {
|
||||
local want="$1" f n
|
||||
for f in "$TEMPLATES_DIR"/*.xml; do
|
||||
[[ -f "$f" ]] || continue
|
||||
n=$(awk 'match($0,/<Name>([^<]+)<\/Name>/,a){print a[1];exit}' "$f")
|
||||
[[ "$n" == "$want" ]] && { echo "$f"; return 0; }
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
remote_has_container() {
|
||||
timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"docker inspect $(printf '%q' "$1") >/dev/null 2>&1" 2>/dev/null
|
||||
}
|
||||
|
||||
remote_state_of() {
|
||||
timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"docker inspect -f '{{.State.Status}}' $(printf '%q' "$1") 2>/dev/null" 2>/dev/null
|
||||
}
|
||||
|
||||
OK=0; FAIL=0; SKIP=0
|
||||
|
||||
case "$MODE" in
|
||||
|
||||
status)
|
||||
# Written as a cache as well as printed. The assistant's fallback_state block cannot afford an
|
||||
# SSH round trip per container mid-question, so it reads this file and reports its AGE — a stale
|
||||
# answer stated as stale is useful, stated as current it is the exact failure this feature
|
||||
# exists to prevent.
|
||||
_present="" _missing=""
|
||||
for c in "${COVERED[@]}"; do
|
||||
if remote_has_container "$c"; then
|
||||
_st=$(remote_state_of "$c")
|
||||
printf ' %-28s on %s (%s)\n' "$c" "$MIRROR" "$_st"
|
||||
_present+="\"$c\":\"${_st:-unknown}\","
|
||||
OK=$((OK+1))
|
||||
else
|
||||
printf ' %-28s MISSING on %s — docker start would fail\n' "$c" "$MIRROR"
|
||||
_missing+="\"$c\","
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
done
|
||||
log "$ICON_FALLBACK Coverage present: $OK · missing: $FAIL"
|
||||
|
||||
mkdir -p "$VV_CACHE_ROOT/api" 2>/dev/null || mkdir -p /tmp/varaverk/api 2>/dev/null
|
||||
_cache="${VV_CACHE_ROOT:-/tmp/varaverk}/api/fallback_presence.json"
|
||||
# Written atomically — a half-written cache read mid-question would report containers as
|
||||
# missing that are merely unparsed.
|
||||
printf '{"present":{%s},"missing":[%s],"partner":"%s","checked":%s}\n' \
|
||||
"${_present%,}" "${_missing%,}" "$MIRROR" "$(date +%s)" > "$_cache.tmp" \
|
||||
&& mv -f "$_cache.tmp" "$_cache"
|
||||
|
||||
[[ "$FAIL" -gt 0 ]] && exit 2 || exit 0
|
||||
;;
|
||||
|
||||
push)
|
||||
# Networks first — a container whose network is absent is created and then cannot start,
|
||||
# which is the failure that read as "auth 0/8, arr 0/5" during onboarding.
|
||||
_nets=()
|
||||
for c in "${COVERED[@]}"; do
|
||||
x=$(xml_for_container "$c") || continue
|
||||
net=$(sed -n 's/.*<Network>\([^<]*\)<\/Network>.*/\1/p' "$x" 2>/dev/null | head -1)
|
||||
net="${net//[[:space:]]/}"
|
||||
# br* is host hardware. wg* is a WireGuard-backed bridge whose meaning does NOT travel:
|
||||
# recreating it on the partner as a plain bridge yields a network that exists, starts its
|
||||
# containers, and routes their traffic OUTSIDE the tunnel. ChannelTube rides wg0 here.
|
||||
case "$net" in
|
||||
''|bridge|host|none|br[0-9]*) continue ;;
|
||||
wg[0-9]*)
|
||||
warn "$c uses $net — a WireGuard-backed network. NOT created on $MIRROR: a plain"
|
||||
warn " bridge of the same name would route its traffic outside the tunnel. Build the"
|
||||
warn " matching tunnel there first, or drop $c from coverage."
|
||||
continue ;;
|
||||
esac
|
||||
_seen=false
|
||||
for n in "${_nets[@]}"; do [[ "$n" == "$net" ]] && { _seen=true; break; }; done
|
||||
[[ "$_seen" == false ]] && _nets+=("$net")
|
||||
done
|
||||
for net in "${_nets[@]}"; do
|
||||
driver=$(timeout "${DOCKER_TIMEOUT:-30}" docker network inspect "$net" --format '{{.Driver}}' 2>/dev/null)
|
||||
if [[ "$driver" != "bridge" ]]; then
|
||||
warn "Network $net is '${driver:-absent}' here, not bridge — create it on $MIRROR by hand"
|
||||
continue
|
||||
fi
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would ensure network $net on $MIRROR"
|
||||
continue
|
||||
fi
|
||||
timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"docker network inspect $(printf '%q' "$net") >/dev/null 2>&1 \
|
||||
|| docker network create --driver bridge $(printf '%q' "$net") >/dev/null" 2>/dev/null \
|
||||
&& log " network $net ready on $MIRROR" \
|
||||
|| warn " could not ensure network $net on $MIRROR"
|
||||
done
|
||||
|
||||
for c in "${COVERED[@]}"; do
|
||||
x=$(xml_for_container "$c") || {
|
||||
warn "$c — no template in $TEMPLATES_DIR names it; skipped"
|
||||
SKIP=$((SKIP+1)); continue
|
||||
}
|
||||
if ! deploy_container_from_xml "$x" "$MIRROR_IP" "$MIRROR_SSH_KEY"; then
|
||||
error "$c — deploy failed"
|
||||
FAIL=$((FAIL+1)); continue
|
||||
fi
|
||||
if [[ "$DRY_RUN" == true ]]; then OK=$((OK+1)); continue; fi
|
||||
|
||||
# Deployed containers must not run here. Stop, then re-inspect — a stop that did not take
|
||||
# is the one outcome that silently duplicates a live service against shared data.
|
||||
timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"docker stop $(printf '%q' "$c") >/dev/null 2>&1" 2>/dev/null
|
||||
st=$(remote_state_of "$c")
|
||||
if [[ "$st" == "running" ]]; then
|
||||
error "$c is RUNNING on $MIRROR after deploy and would not stop — stop it there before continuing"
|
||||
FAIL=$((FAIL+1))
|
||||
else
|
||||
log " $c deployed and ${st:-stopped} on $MIRROR ✅"
|
||||
OK=$((OK+1))
|
||||
fi
|
||||
done
|
||||
log "$ICON_FALLBACK Push complete — deployed $OK · failed $FAIL · skipped $SKIP"
|
||||
[[ "$FAIL" -gt 0 ]] && exit 1 || exit 0
|
||||
;;
|
||||
|
||||
remove)
|
||||
for c in "${COVERED[@]}"; do
|
||||
if ! remote_has_container "$c"; then
|
||||
log " $c not on $MIRROR — nothing to remove"
|
||||
SKIP=$((SKIP+1)); continue
|
||||
fi
|
||||
# Binds are read BEFORE the container goes — once it is removed there is nothing left to
|
||||
# enumerate, and a path list gathered afterwards would silently be empty.
|
||||
_binds=$(timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"docker inspect --format '{{range .HostConfig.Binds}}{{println .}}{{end}}' $(printf '%q' "$c") 2>/dev/null \
|
||||
| awk -F: '{print \$1}'" 2>/dev/null)
|
||||
|
||||
_wipe=()
|
||||
while IFS= read -r _p; do
|
||||
[[ -z "$_p" ]] && continue
|
||||
# Only appdata, and never an appdata root. /mnt/user/appdata as a bind would make one
|
||||
# container removal delete every application on the partner.
|
||||
[[ "$_p" =~ ^/mnt/[^/]+/appdata[^/]*/.+ ]] || {
|
||||
[[ "$_p" =~ ^/mnt/[^/]+/appdata[^/]*/?$ ]] && \
|
||||
warn " $c binds the appdata ROOT ($_p) — refusing to delete it"
|
||||
continue
|
||||
}
|
||||
_wipe+=("$_p")
|
||||
done <<< "$_binds"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would stop and remove $c on $MIRROR"
|
||||
for _p in "${_wipe[@]}"; do warn " DRY RUN — would delete appdata $_p on $MIRROR"; done
|
||||
OK=$((OK+1)); continue
|
||||
fi
|
||||
x=$(xml_for_container "$c") && xml_name=$(basename "$x") || xml_name=""
|
||||
if timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"docker stop $(printf '%q' "$c") >/dev/null 2>&1; \
|
||||
docker rm $(printf '%q' "$c") >/dev/null 2>&1; \
|
||||
${xml_name:+rm -f ${TEMPLATES_DIR}/$(printf '%q' "$xml_name");} \
|
||||
! docker inspect $(printf '%q' "$c") >/dev/null 2>&1" 2>/dev/null; then
|
||||
log " $c removed from $MIRROR ✅"
|
||||
for _p in "${_wipe[@]}"; do
|
||||
if timeout "$SSH_TIMEOUT" ssh -i "$MIRROR_SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=no root@"$MIRROR_IP" \
|
||||
"rm -rf -- $(printf '%q' "$_p") && ! [ -e $(printf '%q' "$_p") ]" 2>/dev/null; then
|
||||
log " appdata deleted on $MIRROR: $_p"
|
||||
else
|
||||
warn " could not delete appdata on $MIRROR: $_p"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
done
|
||||
OK=$((OK+1))
|
||||
else
|
||||
error "$c — removal failed or it still exists on $MIRROR"
|
||||
FAIL=$((FAIL+1))
|
||||
fi
|
||||
done
|
||||
log "$ICON_FALLBACK Remove complete — removed $OK · failed $FAIL · skipped $SKIP"
|
||||
[[ "$FAIL" -gt 0 ]] && exit 1 || exit 0
|
||||
;;
|
||||
esac
|
||||
+321
-56
@@ -6,7 +6,7 @@
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Mutual container fallback between two unRAID servers. Runs continuously as a
|
||||
# background process — started at array start by array_start.sh. Each server
|
||||
# background process — started at array start by array_started.sh. Each server
|
||||
# runs this script independently with no coordination between servers.
|
||||
#
|
||||
# All decisions are based solely on two pings per cycle: remote Tailscale IP
|
||||
@@ -14,7 +14,7 @@
|
||||
# algorithm. Each server acts entirely from its own network perspective.
|
||||
#
|
||||
# Never requires human intervention during normal fallback and handback.
|
||||
# Stop only via User Scripts Abort — do NOT kill directly (state file may corrupt).
|
||||
# Stop only via `fallback.sh --stop` — do NOT kill directly (state file may corrupt).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
@@ -75,10 +75,10 @@
|
||||
# failed handback: containers stop on the covering server, remote goes down
|
||||
# again mid-rsync. Strikes prevent this.
|
||||
#
|
||||
# MY_ID-Based Routing
|
||||
# All tier arrays, DDNS lists, and writeback paths are selected via MY_ID
|
||||
# set by detect_hosts() — not hostname string comparison. The same script
|
||||
# and same config file handle both directions symmetrically.
|
||||
# REMOTE_ID-Based Coverage Config
|
||||
# Tier arrays, writeback paths, and delays are all defined in the covered
|
||||
# host's own conf and read via REMOTE_ID — not MY_ID. Each host owns its
|
||||
# complete recovery profile. The same script handles both directions.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
@@ -89,7 +89,50 @@
|
||||
#
|
||||
# FALLBACK_ENABLED Gate
|
||||
# Exits cleanly when disabled — safe to run on servers being rebuilt without
|
||||
# triggering spurious fallback actions.
|
||||
# triggering spurious fallback actions. Fail-closed: anything that is not exactly
|
||||
# "true" counts as disabled, so a malformed toggle cannot grant this script DDNS
|
||||
# authority and cross-server container control by accident.
|
||||
#
|
||||
# Docker Presence Check
|
||||
# Verifies the docker binary exists before the state machine starts.
|
||||
#
|
||||
# Remote IP Resolution
|
||||
# resolve_remote_ip() must resolve the partner before any SSH operation, so a
|
||||
# remote command can never be issued against an unresolved or stale address.
|
||||
#
|
||||
# Asymmetric Failover / Handback
|
||||
# Entering FALLBACK is immediate — a down partner means users are already affected.
|
||||
# Returning requires FALLBACK_HANDBACK_STRIKES consecutive remote-up checks. The
|
||||
# asymmetry is deliberate: protecting fast costs a few minutes of redundant
|
||||
# coverage, handing back fast on a flapping partner costs a second outage.
|
||||
#
|
||||
# DDNS Excluded From Tier Loops
|
||||
# The Tier 1 stop and start loops explicitly skip any container that is also in
|
||||
# REMOTE_DDNS_CONTAINERS. DDNS is sequenced by the handoff and cutover steps alone,
|
||||
# so ordinary tier processing can never move DNS at the wrong moment.
|
||||
#
|
||||
# Writeback Delay Gate
|
||||
# Tier writeback rsync only runs once the outage has exceeded that tier's writeback
|
||||
# delay. A brief blip does not trigger a full data writeback, which would cost more
|
||||
# than the outage it is compensating for.
|
||||
#
|
||||
# FALLBACK_RSYNC_ENABLED Gate
|
||||
# Writeback is skipped entirely when disabled, and the skip is announced rather than
|
||||
# silent — containers still hand back, but nobody is left assuming data moved.
|
||||
#
|
||||
# Play State Sync Before Cutover
|
||||
# Handback retries play_state_sync up to PLAY_SYNC_HANDBACK_RETRIES times before DNS
|
||||
# cuts over, so users land on current watch state. Exhausting retries warns and
|
||||
# proceeds — stale resume positions are not worth holding DNS on a downed service.
|
||||
#
|
||||
# Partnership Suspend Abort
|
||||
# If partnership goes inactive mid-fallback, _abort_fallback_containers() stops the
|
||||
# fallback containers and returns to NORMAL rather than leaving this host serving a
|
||||
# partner it is no longer paired with.
|
||||
#
|
||||
# Container Verify Wait
|
||||
# CONTAINER_VERIFY_WAIT seconds elapse after each start before the running check, so
|
||||
# a container that starts and immediately crashes is caught rather than counted up.
|
||||
#
|
||||
# Version Parity Check
|
||||
# Refuses handback if remote unRAID version doesn't match. A mismatch may
|
||||
@@ -115,6 +158,13 @@
|
||||
# detect_hosts() resolves MY_ID / REMOTE_ID from master.conf at startup.
|
||||
# Exits if the host cannot be identified — prevents running on an unknown machine.
|
||||
#
|
||||
# Partnership Gate
|
||||
# Fallback is only permitted when the local partnership DB reports an active
|
||||
# partnership. If partnership goes inactive, a grace timer starts. After
|
||||
# FALLBACK_PARTNERSHIP_SUSPEND_AFTER minutes the loop suspends: no new FALLBACK
|
||||
# is entered, and any live fallback containers are stopped cleanly. Resumes
|
||||
# automatically when partnership becomes active again.
|
||||
#
|
||||
# Silent by Default
|
||||
# State transitions: warn() — always visible. Healthy routine cycles: log()
|
||||
# — suppressed unless --log. Produces no output on clean cycles.
|
||||
@@ -123,10 +173,13 @@
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# FALLBACK_STATE_FILE — /boot/config/fallback_state.db (survives reboots)
|
||||
# FALLBACK_STATE_FILE — $STATE_DIR/fallback_state.db (survives reboots)
|
||||
# Keys: state, fallback_start, handback_strikes, tier2_started,
|
||||
# tier3_started, tier4_started. Lives on /boot/ intentionally — if the
|
||||
# server was in FALLBACK when it rebooted, it resumes FALLBACK on restart.
|
||||
# tier3_started, tier4_started, partnership_suspended, partner_lost_at.
|
||||
# Survives reboots — $STATE_DIR is on the boot device (internal) or appdata
|
||||
# (flash). Either way the array is up before this script runs, so the file
|
||||
# is always accessible. If the server was in FALLBACK when it rebooted,
|
||||
# it resumes FALLBACK on restart.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -141,9 +194,10 @@
|
||||
# FALLBACK_HOST*_STOP_ON_NO_NET
|
||||
# Containers stopped when this host loses internet
|
||||
#
|
||||
# FALLBACK_HOST*_COVERS_HOST*_TIER1–4
|
||||
# Containers this host starts for the remote when remote is down, by tier.
|
||||
# Variable pattern: FALLBACK_${MY_ID}_COVERS_${REMOTE_ID}_TIER${N}
|
||||
# FALLBACK_HOST*_TIER1–4
|
||||
# Containers started on this host when the partner is down — defined in the
|
||||
# partner's own conf (what the partner wants covered). Read via REMOTE_ID.
|
||||
# Variable pattern: FALLBACK_${REMOTE_ID}_TIER${N}
|
||||
#
|
||||
# HOST*_TIER2_DELAY / TIER3_DELAY / TIER4_DELAY
|
||||
# Minutes after FALLBACK entry before activating each tier.
|
||||
@@ -164,6 +218,11 @@
|
||||
# FALLBACK_ENABLED
|
||||
# false = exit cleanly (e.g. remote server being rebuilt). (default: false)
|
||||
#
|
||||
# FALLBACK_PARTNERSHIP_SUSPEND_AFTER
|
||||
# Minutes without an active partnership before fallback suspends itself.
|
||||
# Grace period: fallback continues normally during this window.
|
||||
# 0 = suspend immediately when partnership goes inactive. (default: 120)
|
||||
#
|
||||
# FALLBACK_CHECK_INTERVAL
|
||||
# Seconds between connectivity checks. (default: 30)
|
||||
#
|
||||
@@ -171,7 +230,7 @@
|
||||
# Consecutive remote-up checks required before handback begins. (default: 3)
|
||||
#
|
||||
# FALLBACK_STATE_FILE
|
||||
# State file path — /boot/config/fallback_state.db — survives reboots.
|
||||
# State file path — $STATE_DIR/fallback_state.db — survives reboots.
|
||||
#
|
||||
# FALLBACK_RSYNC_ENABLED
|
||||
# Gate for writeback rsync jobs during handback. (default: true)
|
||||
@@ -184,7 +243,7 @@
|
||||
# ==============================================================================================
|
||||
#
|
||||
# fallback.sh
|
||||
# Normal start — continuous loop. Start via User Scripts or array_start.sh.
|
||||
# Normal start — continuous loop. Started by array_started.sh at array start.
|
||||
#
|
||||
# fallback.sh --stop
|
||||
# Gracefully stop the running instance (SIGTERM → wait 10s → SIGKILL).
|
||||
@@ -200,8 +259,7 @@
|
||||
# fallback.sh --log
|
||||
# Verbose output on every decision in every cycle.
|
||||
#
|
||||
# To stop: use `fallback.sh --stop` or click Abort in User Scripts.
|
||||
# Do NOT kill -9 directly — state file may corrupt if mid-write.
|
||||
# To stop: use `fallback.sh --stop`. Do NOT kill -9 — state file may corrupt if mid-write.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
@@ -256,8 +314,55 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# FALLBACK_ENABLED gate — exits cleanly when disabled (e.g. HOST2 being rebuilt)
|
||||
if [[ "${FALLBACK_ENABLED:-false}" == false ]]; then
|
||||
|
||||
# ── Persistent dry-run ────────────────────────────────────────────────────────────────────────
|
||||
# Conf-driven, not argument-driven, deliberately. array_started.sh launches every entry as a bare
|
||||
# `bash script.sh &` with no arguments, so a --dry-run typed at a terminal survives exactly until
|
||||
# the next array start — and then the LIVE daemon comes up in its place, silently, which is the
|
||||
# one transition nobody would be watching for.
|
||||
#
|
||||
# Setting it here means the mode is a property of the install rather than of how the process
|
||||
# happened to be started: array start, the Fallback tab's button, and a hand-run all agree.
|
||||
#
|
||||
# OR, never override: --dry-run on the command line still wins over a conf that says false, so an
|
||||
# ad-hoc preview against a live install needs no conf edit.
|
||||
if [[ "${FALLBACK_DRY_RUN:-false}" == "true" ]]; then
|
||||
DRY_RUN=true
|
||||
fi
|
||||
|
||||
# ── Persistent log ────────────────────────────────────────────────────────────────────────────
|
||||
# array_started.sh launches every entry as a bare `bash script.sh &` with no redirection, so this
|
||||
# daemon's output has never been captured anywhere: /var/log/varaverk has a directory for every
|
||||
# other script family and none for Fallback. A month of dry-run observation would have persisted
|
||||
# nothing at all.
|
||||
#
|
||||
# /var/log is a 128 MB tmpfs on Unraid — RAM, and cleared on reboot — so the log goes to
|
||||
# LOG_ARCHIVE_DIR, which follows DATA_DIR onto real storage.
|
||||
#
|
||||
# Only when stdout is not a terminal. Run by hand, output still goes to the terminal exactly as
|
||||
# before; run by array_started or the Fallback tab's button, it lands in the file. A plain append
|
||||
# redirect rather than `tee` through process substitution: no extra child to outlive, and nothing
|
||||
# for the shutdown trap to race.
|
||||
FALLBACK_LOG="${LOG_ARCHIVE_DIR:-${DATA_DIR:-/tmp}/logs}/fallback.log"
|
||||
if [[ ! -t 1 ]]; then
|
||||
mkdir -p "$(dirname "$FALLBACK_LOG")" 2>/dev/null
|
||||
# One rotation, sized rather than line-counted — the whole point of this log is a long run,
|
||||
# and _orch_trim_log()'s 1000-line cap would discard weeks of it. Event-only output (no
|
||||
# --log) is a few lines per incident, so this holds years; --log fills it in about a fortnight
|
||||
# and then keeps the most recent fortnight plus the one before it.
|
||||
_fb_max=$(( ${FALLBACK_LOG_MAX_MB:-5} * 1048576 ))
|
||||
if [[ -f "$FALLBACK_LOG" ]] && (( $(stat -c %s "$FALLBACK_LOG" 2>/dev/null || echo 0) > _fb_max )); then
|
||||
mv -f "$FALLBACK_LOG" "${FALLBACK_LOG}.1" 2>/dev/null
|
||||
fi
|
||||
exec >> "$FALLBACK_LOG" 2>&1
|
||||
echo ""
|
||||
echo "═══ fallback.sh started $(date '+%Y-%m-%d %H:%M:%S') — dry_run=${DRY_RUN} pid=$$ ═══"
|
||||
fi
|
||||
# FALLBACK_ENABLED gate — exits cleanly when disabled.
|
||||
# Fail-closed: anything that isn't exactly "true" disables fallback. Matching only the
|
||||
# literal "false" would let a typo ("no", "0", "FALSE") hand this script DDNS authority
|
||||
# and cross-server container control on a toggle nobody meant to set.
|
||||
if [[ "${FALLBACK_ENABLED:-false}" != "true" ]]; then
|
||||
warn "FALLBACK_ENABLED=false — fallback monitoring disabled"
|
||||
warn "Set FALLBACK_ENABLED=true in master.conf when both servers are ready"
|
||||
exit 0
|
||||
@@ -271,14 +376,9 @@ if ! command -v docker &>/dev/null; then
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
require_partnership
|
||||
resolve_remote_ip
|
||||
|
||||
# Validate unRAID notify script — used throughout for state change notifications
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no container or DDNS changes will be made"
|
||||
|
||||
# Timeout for all docker and SSH docker commands
|
||||
@@ -286,10 +386,13 @@ DOCKER_TIMEOUT=15
|
||||
SSH_TIMEOUT=10
|
||||
CONTAINER_VERIFY_WAIT=5 # seconds after start before verifying container is up
|
||||
|
||||
log "$ICON_GEAR Config: check-interval=${FALLBACK_CHECK_INTERVAL}s handback-strikes=${FALLBACK_HANDBACK_STRIKES} rsync-enabled=${FALLBACK_RSYNC_ENABLED:-true}"
|
||||
log "$ICON_GEAR Timeouts: docker=${DOCKER_TIMEOUT}s ssh=${SSH_TIMEOUT}s verify-wait=${CONTAINER_VERIFY_WAIT}s"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── STATE FILE HELPERS ────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# State file on /boot/config — survives reboots.
|
||||
# State file at $FALLBACK_STATE_FILE ($STATE_DIR/fallback_state.db) — survives reboots.
|
||||
# Format: key=value one per line.
|
||||
# Keys: state, fallback_start, handback_strikes, tier2_started, tier3_started, tier4_started
|
||||
|
||||
@@ -300,21 +403,69 @@ state_get() {
|
||||
state_set() {
|
||||
local key="$1" value="$2"
|
||||
if grep -q "^${key}=" "$FALLBACK_STATE_FILE" 2>/dev/null; then
|
||||
sed -i "s|^${key}=.*|${key}=${value}|" "$FALLBACK_STATE_FILE"
|
||||
# Escape sed replacement metacharacters: \ first, then & and |
|
||||
local safe="${value//\\/\\\\}"; safe="${safe//&/\\&}"; safe="${safe//|/\\|}"
|
||||
sed -i "s|^${key}=.*|${key}=${safe}|" "$FALLBACK_STATE_FILE"
|
||||
else
|
||||
echo "${key}=${value}" >> "$FALLBACK_STATE_FILE"
|
||||
fi
|
||||
}
|
||||
|
||||
state_init() {
|
||||
# A dry run must not leave the host believing it failed over. state_set() writes
|
||||
# unconditionally, and this file survives reboots and is what the real daemon — and the
|
||||
# Monitor and Fallback cards — read to decide what is happening. A --dry-run walk through
|
||||
# FAILOVER would have written state=FALLBACK, the tier flags and the strike counter into it
|
||||
# for real, and nothing would have put them back.
|
||||
#
|
||||
# Copied rather than merely redirected, so the preview still starts from the live state and
|
||||
# can advance through tiers exactly as a real run would. The copy lands in the RAM cache and
|
||||
# dies with the reboot.
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
local live="$FALLBACK_STATE_FILE"
|
||||
FALLBACK_STATE_FILE="${VV_CACHE_ROOT:-/tmp/varaverk}/fallback_state.dryrun.$$"
|
||||
mkdir -p "$(dirname "$FALLBACK_STATE_FILE")"
|
||||
if [[ -f "$live" ]]; then cp -f "$live" "$FALLBACK_STATE_FILE"; else : > "$FALLBACK_STATE_FILE"; fi
|
||||
warn "DRY RUN — state writes redirected to $FALLBACK_STATE_FILE (live state untouched)"
|
||||
fi
|
||||
mkdir -p "$(dirname "$FALLBACK_STATE_FILE")"
|
||||
[[ ! -f "$FALLBACK_STATE_FILE" ]] && touch "$FALLBACK_STATE_FILE"
|
||||
[[ -z "$(state_get state)" ]] && state_set state "NORMAL"
|
||||
[[ -z "$(state_get fallback_start)" ]] && state_set fallback_start "0"
|
||||
[[ -z "$(state_get handback_strikes)" ]] && state_set handback_strikes "0"
|
||||
[[ -z "$(state_get tier2_started)" ]] && state_set tier2_started "false"
|
||||
[[ -z "$(state_get tier3_started)" ]] && state_set tier3_started "false"
|
||||
[[ -z "$(state_get tier4_started)" ]] && state_set tier4_started "false"
|
||||
[[ -z "$(state_get state)" ]] && state_set state "NORMAL"
|
||||
[[ -z "$(state_get fallback_start)" ]] && state_set fallback_start "0"
|
||||
[[ -z "$(state_get handback_strikes)" ]] && state_set handback_strikes "0"
|
||||
[[ -z "$(state_get tier2_started)" ]] && state_set tier2_started "false"
|
||||
[[ -z "$(state_get tier3_started)" ]] && state_set tier3_started "false"
|
||||
[[ -z "$(state_get tier4_started)" ]] && state_set tier4_started "false"
|
||||
[[ -z "$(state_get partnership_suspended)" ]] && state_set partnership_suspended "false"
|
||||
[[ -z "$(state_get partner_lost_at)" ]] && state_set partner_lost_at "0"
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── PARTNERSHIP GATE HELPERS ──────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Returns 0 (true) if the local partnership DB reports an ACTIVE partnership.
|
||||
# Reads /boot/config/partnership_<hostname>.db — no subprocess, no SSH.
|
||||
check_partnership_active() {
|
||||
local state_file="${STATE_DIR}/partnership_${LOCAL_SERVER_NAME}.db"
|
||||
local state
|
||||
state=$(grep "^state=" "$state_file" 2>/dev/null | cut -d= -f2)
|
||||
[[ "$state" == "ACTIVE" ]]
|
||||
}
|
||||
|
||||
# Stop all tiers that were started during a fallback and clear their state flags.
|
||||
# Called when the partnership gate suspends an in-progress fallback.
|
||||
_abort_fallback_containers() {
|
||||
for tier in 4 3 2 1; do
|
||||
[[ "$(state_get "tier${tier}_started")" != "true" ]] && continue
|
||||
warn "Aborting Tier $tier containers — partnership suspended"
|
||||
local containers
|
||||
read -r -a containers <<< "$(get_tier_containers "$tier")"
|
||||
for container in "${containers[@]}"; do
|
||||
[[ -n "$container" ]] && local_stop "$container"
|
||||
done
|
||||
state_set "tier${tier}_started" "false"
|
||||
done
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -322,7 +473,7 @@ state_init() {
|
||||
# ==============================================================================================
|
||||
# All docker commands wrapped in DOCKER_TIMEOUT.
|
||||
# All SSH docker commands wrapped in SSH_TIMEOUT.
|
||||
# Verification after start — failover is critical, confirm containers came up.
|
||||
# Verification after start — fallback is critical, confirm containers came up.
|
||||
|
||||
# Start a container locally and verify it came up
|
||||
local_start() {
|
||||
@@ -346,7 +497,7 @@ local_start() {
|
||||
post_status=$(timeout "$DOCKER_TIMEOUT" docker inspect -f '{{.State.Running}}' \
|
||||
"$container" 2>/dev/null)
|
||||
if [[ "$post_status" == "true" ]]; then
|
||||
log "$ICON_STARTED $container started and running ✅"
|
||||
echo "$ICON_STARTED $container started and running ✅"
|
||||
return 0
|
||||
else
|
||||
error "$container started but crashed immediately"
|
||||
@@ -379,7 +530,7 @@ local_stop() {
|
||||
return 0
|
||||
fi
|
||||
timeout "$DOCKER_TIMEOUT" docker stop "$container" >/dev/null 2>&1 && \
|
||||
log "$ICON_STOPPED $container stopped" || \
|
||||
echo "$ICON_STOPPED $container stopped" || \
|
||||
error "Failed to stop $container locally"
|
||||
}
|
||||
|
||||
@@ -411,7 +562,7 @@ remote_start() {
|
||||
"timeout $DOCKER_TIMEOUT docker inspect -f '{{.State.Running}}' \
|
||||
$container 2>/dev/null" 2>/dev/null)
|
||||
if [[ "$post_status" == "true" ]]; then
|
||||
log "$ICON_STARTED $container started on $REMOTE_SERVER_NAME ✅"
|
||||
echo "$ICON_STARTED $container started on $REMOTE_SERVER_NAME ✅"
|
||||
return 0
|
||||
else
|
||||
error "$container started on $REMOTE_SERVER_NAME but crashed immediately"
|
||||
@@ -448,7 +599,7 @@ remote_stop() {
|
||||
timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" -o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
root@"$REMOTE_SERVER" \
|
||||
"timeout $DOCKER_TIMEOUT docker stop $container" >/dev/null 2>&1 && \
|
||||
log "$ICON_STOPPED $container stopped on $REMOTE_SERVER_NAME" || \
|
||||
echo "$ICON_STOPPED $container stopped on $REMOTE_SERVER_NAME" || \
|
||||
error "Failed to stop $container on $REMOTE_SERVER_NAME"
|
||||
}
|
||||
|
||||
@@ -488,8 +639,7 @@ remote_ddns_stop() {
|
||||
|
||||
get_tier_containers() {
|
||||
local tier="$1"
|
||||
local remote_id="${REMOTE_ID}"
|
||||
local var_name="FALLBACK_${MY_ID}_COVERS_${remote_id}_TIER${tier}"
|
||||
local var_name="FALLBACK_${REMOTE_ID}_TIER${tier}"
|
||||
eval "echo \"\${${var_name}[@]:-}\""
|
||||
}
|
||||
|
||||
@@ -554,13 +704,13 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
TIER4=$(state_get tier4_started)
|
||||
STRIKES=$(state_get handback_strikes)
|
||||
|
||||
local_ver=$(grep -oP '(?<=version=")[^"]+' /etc/unraid-version 2>/dev/null || echo "unknown")
|
||||
local_ver=$(platform_get_os_version 2>/dev/null || echo "unknown")
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY FALLBACK STATUS ━━━━━"
|
||||
echo "$ICON_HOST My ID: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_HOST Remote ID: $REMOTE_ID ($REMOTE_SERVER_NAME — $REMOTE_SERVER)"
|
||||
echo "$ICON_GEAR unRAID ver: $local_ver"
|
||||
echo "$ICON_GEAR OS ver: $local_ver"
|
||||
echo "$ICON_FALLBACK State: $CURRENT_STATE"
|
||||
echo "$ICON_NET Local DDNS: ${LOCAL_DDNS_CONTAINERS[*]:-none}"
|
||||
echo "$ICON_NET Remote DDNS: ${REMOTE_DDNS_CONTAINERS[*]:-none}"
|
||||
@@ -577,6 +727,18 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
|
||||
echo "$ICON_GEAR Enabled: $FALLBACK_ENABLED"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
|
||||
SUSPENDED=$(state_get partnership_suspended)
|
||||
PLOST=$(state_get partner_lost_at)
|
||||
if [[ "$SUSPENDED" == "true" ]]; then
|
||||
echo "$ICON_GEAR Partnership: SUSPENDED (no active partnership)"
|
||||
elif [[ -n "$PLOST" && "$PLOST" != "0" ]]; then
|
||||
PLOST_MIN=$(( ($(date +%s) - PLOST) / 60 ))
|
||||
echo "$ICON_GEAR Partnership: GRACE (${PLOST_MIN}min / ${FALLBACK_PARTNERSHIP_SUSPEND_AFTER:-120}min until suspend)"
|
||||
else
|
||||
echo "$ICON_GEAR Partnership: ACTIVE"
|
||||
fi
|
||||
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
@@ -597,7 +759,7 @@ run_handback() {
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Pre-flight ━━━"
|
||||
|
||||
if ! check_unraid_version_parity; then
|
||||
if ! check_os_version_parity; then
|
||||
warn "Version parity check failed — aborting handback, will retry next cycle"
|
||||
state_set handback_strikes 0
|
||||
return 1
|
||||
@@ -657,8 +819,16 @@ run_handback() {
|
||||
for job in "${jobs[@]}"; do
|
||||
[[ -z "$job" ]] && continue
|
||||
log "Syncing: $job"
|
||||
[[ "$DRY_RUN" == false ]] && bash "$SCRIPT_DIR/../Rsync/rsync.sh" "$job" \
|
||||
|| warn "DRY RUN — would rsync: $job"
|
||||
# if/else, not A && B || C. In the shorthand a REAL run whose rsync exits
|
||||
# non-zero falls through to the || branch and logs "DRY RUN — would rsync",
|
||||
# so a failed Tier writeback reported itself as a preview and the real
|
||||
# failure went unsaid. The Tier 1 block below always had this right.
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
bash "$SCRIPT_DIR/../Rsync/rsync.sh" "$job" \
|
||||
|| error "Tier $tier writeback FAILED: $job"
|
||||
else
|
||||
warn "DRY RUN — would rsync: $job"
|
||||
fi
|
||||
done
|
||||
else
|
||||
log "Tier $tier writeback skipped — outage ${outage_minutes}min < ${threshold}min"
|
||||
@@ -714,11 +884,7 @@ run_handback() {
|
||||
[[ -z "$job" ]] && continue
|
||||
log "Syncing: $job"
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
if [[ "$(basename "$job")" == "Emby" ]]; then
|
||||
bash "$SCRIPT_DIR/../Rsync/rsync.sh" "$job" --profile=emby-fallback
|
||||
else
|
||||
bash "$SCRIPT_DIR/../Rsync/rsync.sh" "$job"
|
||||
fi
|
||||
bash "$SCRIPT_DIR/../Rsync/rsync.sh" "$job"
|
||||
else
|
||||
warn "DRY RUN — would rsync: $job"
|
||||
fi
|
||||
@@ -739,7 +905,33 @@ run_handback() {
|
||||
[[ "$is_ddns" == false ]] && remote_start "$container"
|
||||
done
|
||||
|
||||
# ── Step 7: DNS Cutover — final step ────────────────────────────────────────────────────
|
||||
# ── Step 7: Play State Sync — DNS held until sync succeeds or retries exhausted ──────────
|
||||
# Retried so users land on current watch state — after a 3hr outage one successful
|
||||
# run catches everything regardless of how many cron cycles were missed.
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Play State Sync ━━━"
|
||||
local _sync_ok=false
|
||||
local _retry_max="${PLAY_SYNC_HANDBACK_RETRIES:-5}"
|
||||
local _retry_delay="${PLAY_SYNC_HANDBACK_RETRY_DELAY:-60}"
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
for _attempt in $(seq 1 "$_retry_max"); do
|
||||
if bash "$SCRIPT_DIR/../Media/play_state_sync.sh" --wait; then
|
||||
_sync_ok=true
|
||||
log "Play state synced ✅ (attempt ${_attempt}/${_retry_max})"
|
||||
break
|
||||
fi
|
||||
if [[ "$_attempt" -lt "$_retry_max" ]]; then
|
||||
warn "Play state sync failed (attempt ${_attempt}/${_retry_max}) — retrying in ${_retry_delay}s"
|
||||
sleep "$_retry_delay"
|
||||
fi
|
||||
done
|
||||
[[ "$_sync_ok" == false ]] && \
|
||||
warn "Play state sync failed after ${_retry_max} attempts — proceeding to DNS cutover"
|
||||
else
|
||||
warn "DRY RUN — would retry play_state_sync --wait up to ${_retry_max} times before DNS cutover"
|
||||
fi
|
||||
|
||||
# ── Step 8: DNS Cutover — final step ────────────────────────────────────────────────────
|
||||
# DNS cuts over ONLY after Tier 1 containers confirmed up
|
||||
echo ""
|
||||
echo "━━━ $ICON_NET DNS Cutover ━━━"
|
||||
@@ -749,7 +941,7 @@ run_handback() {
|
||||
done
|
||||
warn "$ICON_NET Remote DDNS started — DNS now points at $REMOTE_SERVER_NAME"
|
||||
|
||||
# ── Step 8: Return to NORMAL ─────────────────────────────────────────────────────────────
|
||||
# ── Step 9: Return to NORMAL ─────────────────────────────────────────────────────────────
|
||||
echo ""
|
||||
state_set state "NORMAL"
|
||||
state_set fallback_start "0"
|
||||
@@ -780,6 +972,19 @@ echo " $ICON_NET Remote IP: $REMOTE_SERVER"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
FALLBACK_RUNNING=true
|
||||
|
||||
# The dry-run state copy is per-PID and would otherwise accumulate one file per preview run.
|
||||
# EXIT as well as the signals, because a dry run is usually ended with Ctrl-C or --stop but can
|
||||
# also just fall out of the loop.
|
||||
dryrun_state_cleanup() {
|
||||
[[ "$DRY_RUN" == true && "$FALLBACK_STATE_FILE" == *".dryrun."* ]] && rm -f "$FALLBACK_STATE_FILE"
|
||||
return 0
|
||||
}
|
||||
# Must call _release_all_locks too. acquire_lock() registers its own EXIT trap, and bash keeps
|
||||
# exactly one per signal — a bare `trap ... EXIT` here silently replaced it and orphaned
|
||||
# fallback.lock, which is the precise failure the _LOCK_FILES registry in common.sh was built to
|
||||
# stop. The signal trap only needs `exit 0`; that fires EXIT, which does both jobs.
|
||||
trap 'dryrun_state_cleanup; _release_all_locks' EXIT
|
||||
trap 'FALLBACK_RUNNING=false; warn "Fallback received shutdown signal — stopping cleanly"; exit 0' \
|
||||
SIGTERM SIGINT
|
||||
|
||||
@@ -796,6 +1001,59 @@ while [[ "$FALLBACK_RUNNING" == true ]]; do
|
||||
|
||||
log "Remote: $REMOTE_UP | Internet: $INTERNET_UP | State: $CURRENT_STATE"
|
||||
|
||||
# ── Partnership gate ──────────────────────────────────────────────────────────────────────
|
||||
# ── Partnership Gate ──────────────────────────────────────────────────────────────────────
|
||||
# A fallback without an active partner is meaningless — who are we covering for?
|
||||
# The gate suspends the loop after a grace period and aborts any live fallback
|
||||
# containers. Clears automatically when partnership becomes active again.
|
||||
if check_partnership_active; then
|
||||
# Partnership healthy — clear any accumulated grace state
|
||||
if [[ "$(state_get partnership_suspended)" == "true" ]]; then
|
||||
warn "Active partnership restored — re-enabling fallback monitoring"
|
||||
state_set partnership_suspended "false"
|
||||
state_set partner_lost_at "0"
|
||||
notify "Fallback re-enabled on $(hostname) — active partnership restored" \
|
||||
"Fallback" "normal"
|
||||
elif [[ "$(state_get partner_lost_at)" != "0" ]]; then
|
||||
log "Partnership recovered within grace window — resetting timer"
|
||||
state_set partner_lost_at "0"
|
||||
fi
|
||||
else
|
||||
PARTNER_LOST_AT=$(state_get partner_lost_at)
|
||||
if [[ -z "$PARTNER_LOST_AT" || "$PARTNER_LOST_AT" == "0" ]]; then
|
||||
warn "No active partnership detected — grace timer started (${FALLBACK_PARTNERSHIP_SUSPEND_AFTER:-120}min before suspend)"
|
||||
state_set partner_lost_at "$NOW"
|
||||
PARTNER_LOST_AT=$NOW
|
||||
fi
|
||||
|
||||
PARTNER_LOST_MIN=$(( (NOW - PARTNER_LOST_AT) / 60 ))
|
||||
SUSPEND_AFTER=${FALLBACK_PARTNERSHIP_SUSPEND_AFTER:-120}
|
||||
|
||||
if [[ "$(state_get partnership_suspended)" == "true" || "$PARTNER_LOST_MIN" -ge "$SUSPEND_AFTER" ]]; then
|
||||
# Suspend — abort any live fallback containers, then skip this cycle entirely
|
||||
if [[ "$(state_get partnership_suspended)" != "true" ]]; then
|
||||
warn "No active partnership for ${PARTNER_LOST_MIN}min — suspending fallback"
|
||||
state_set partnership_suspended "true"
|
||||
notify "Fallback suspended on $(hostname) — no active partnership for ${PARTNER_LOST_MIN}min" \
|
||||
"Fallback" "warning"
|
||||
fi
|
||||
if [[ "$CURRENT_STATE" == "FALLBACK" ]]; then
|
||||
warn "Aborting in-progress fallback — partnership suspended"
|
||||
_abort_fallback_containers
|
||||
state_set state "NORMAL"
|
||||
state_set fallback_start "0"
|
||||
state_set handback_strikes "0"
|
||||
notify "Fallback aborted on $(hostname) — no active partnership, containers stopped" \
|
||||
"Fallback" "warning"
|
||||
fi
|
||||
log "Partnership suspended — skipping fallback cycle"
|
||||
sleep "$FALLBACK_CHECK_INTERVAL" & wait $!
|
||||
continue
|
||||
else
|
||||
log "No active partnership — ${PARTNER_LOST_MIN}min elapsed / ${SUSPEND_AFTER}min until suspend"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ════════════════════════════════════════════════════════════════
|
||||
# NORMAL STATE
|
||||
# ════════════════════════════════════════════════════════════════
|
||||
@@ -867,17 +1125,24 @@ while [[ "$FALLBACK_RUNNING" == true ]]; do
|
||||
fi
|
||||
|
||||
elif [[ "$INTERNET_UP" == false ]]; then
|
||||
# Lost internet during fallback — enter DARK
|
||||
# Lost internet during fallback — enter DARK — same actions as NO_INTERNET
|
||||
echo ""
|
||||
echo "━━━ $ICON_NET Entering DARK — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
warn "Lost internet during fallback — entering DARK state"
|
||||
state_set state "DARK"
|
||||
local_ddns_stop
|
||||
|
||||
# Stop containers configured to stop on internet loss
|
||||
read -r -a stop_on_no_net <<< "$(get_stop_on_no_net)"
|
||||
for container in "${stop_on_no_net[@]}"; do
|
||||
[[ -n "$container" ]] && local_stop "$container"
|
||||
done
|
||||
|
||||
notify "DARK state on $(hostname) — lost internet during fallback" \
|
||||
"Fallback" "warning"
|
||||
|
||||
else
|
||||
# Still in failover — check tier escalation
|
||||
# Still in FALLBACK state — check tier escalation
|
||||
state_set handback_strikes "0"
|
||||
|
||||
TIER2_DELAY=$(get_tier2_delay)
|
||||
@@ -973,7 +1238,7 @@ while [[ "$FALLBACK_RUNNING" == true ]]; do
|
||||
|
||||
if [[ "$REMOTE_UP" == true ]]; then
|
||||
warn "Remote up, internet up — transitioning through FALLBACK for handback"
|
||||
# Was in failover before DARK — route through FALLBACK state for handback
|
||||
# Was in FALLBACK state before DARK — route through FALLBACK state for handback
|
||||
state_set state "FALLBACK"
|
||||
state_set handback_strikes "0"
|
||||
else
|
||||
@@ -1001,4 +1266,4 @@ while [[ "$FALLBACK_RUNNING" == true ]]; do
|
||||
sleep "$FALLBACK_CHECK_INTERVAL" &
|
||||
wait $!
|
||||
|
||||
done
|
||||
done
|
||||
|
||||
Regular → Executable
+141
-24
@@ -41,9 +41,23 @@
|
||||
# ==============================================================================================
|
||||
#
|
||||
# iptables Safety Trap
|
||||
# The DROP rule is removed via trap on ANY exit — normal completion, crash, error,
|
||||
# ctrl-c. Remote connectivity is always restored regardless of test outcome.
|
||||
# You cannot accidentally leave the remote permanently blocked.
|
||||
# The DROP rule is removed via an EXIT trap that fires on normal completion, error
|
||||
# exit, script crash, ctrl-c (SIGINT) and SIGTERM — verified, not assumed. Remote
|
||||
# connectivity is restored regardless of test outcome.
|
||||
#
|
||||
# Stale Rule Sweep
|
||||
# The trap above cannot cover SIGKILL or a power cut, which are the only ways a DROP
|
||||
# rule survives the test. One stranded that way makes fallback.sh see the partner as
|
||||
# permanently down and hold FALLBACK indefinitely, so pre-flight clears any leftover
|
||||
# rule before doing anything else — including before the reachability check, which
|
||||
# would otherwise fail and blame the network for the test's own residue.
|
||||
#
|
||||
# Root Enforcement
|
||||
# iptables and container control require root.
|
||||
#
|
||||
# iptables Presence Check
|
||||
# platform_require_cmd confirms iptables exists before the test begins — there is no
|
||||
# point entering a connectivity simulation that cannot simulate anything.
|
||||
#
|
||||
# FALLBACK_ENABLED Gate
|
||||
# Aborts if FALLBACK_ENABLED=false. Testing a disabled fallback system is
|
||||
@@ -77,13 +91,13 @@
|
||||
# FALLBACK_TEST_BLOCK_WAIT
|
||||
# Seconds to wait in Phase 3 for fallback.sh to detect the outage.
|
||||
# Must be > FALLBACK_CHECK_INTERVAL + buffer. At 30s interval: use ≥60s.
|
||||
# (default: 60)
|
||||
# (shipped default: 150)
|
||||
#
|
||||
# FALLBACK_TEST_HANDBACK_WAIT
|
||||
# Seconds to wait in Phase 6 for fallback.sh to complete handback.
|
||||
# Must cover: FALLBACK_HANDBACK_STRIKES × FALLBACK_CHECK_INTERVAL + rsync
|
||||
# duration + container start time. At 3 strikes × 30s + ~2min rsync +
|
||||
# ~1min container start: use ≥240s. (default: 300)
|
||||
# ~1min container start: use ≥240s. (shipped default: 360)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
@@ -104,6 +118,11 @@
|
||||
# fallback_test.sh --log
|
||||
# Verbose output on every check in every phase.
|
||||
#
|
||||
# fallback_test.sh --stop
|
||||
# Stop a running test. SIGTERM only — never SIGKILL, because only this script's EXIT
|
||||
# trap removes the iptables DROP rule it installed. Also sweeps a rule stranded by an
|
||||
# earlier SIGKILL or power cut.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
@@ -112,6 +131,58 @@ source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ── Stop mode — runs before acquire_lock so we can target the holding instance ────────────────
|
||||
#
|
||||
# SIGTERM ONLY, and deliberately no SIGKILL escalation — the opposite of fallback.sh --stop.
|
||||
# A running test holds an iptables DROP rule against the partner, and the only thing that removes
|
||||
# it is this script's own EXIT trap. SIGKILL does not run traps, so force-killing a test strands
|
||||
# the rule: the partner stays invisible, fallback.sh reads that as a permanent outage and holds
|
||||
# FALLBACK indefinitely. A test that will not die is a worse outcome than a test still running,
|
||||
# so this reports the stranded rule and the command to clear it rather than causing one.
|
||||
if [[ " ${PARSED_ARGS[*]:-} " == *" --stop "* ]]; then
|
||||
LOCKFILE="${LOCK_DIR}/fallback_test.lock"
|
||||
if [[ ! -f "$LOCKFILE" ]]; then
|
||||
log "No fallback_test.sh lock found — not running"
|
||||
exit 0
|
||||
fi
|
||||
lock_content=$(cat "$LOCKFILE" 2>/dev/null)
|
||||
target_pid="${lock_content%%:*}"
|
||||
if [[ -z "$target_pid" ]] || ! kill -0 "$target_pid" 2>/dev/null; then
|
||||
warn "Stale lock — fallback_test.sh not running (PID ${target_pid:-unknown} gone) — clearing"
|
||||
rm -f "$LOCKFILE"
|
||||
# A stale lock is exactly the SIGKILL/power-cut case, so the rule may still be in place.
|
||||
if iptables -C OUTPUT -d "${REMOTE_SERVER:-0.0.0.0}" -j DROP 2>/dev/null; then
|
||||
warn "Stranded iptables DROP rule found for $REMOTE_SERVER — removing"
|
||||
iptables -D OUTPUT -d "$REMOTE_SERVER" -j DROP 2>/dev/null \
|
||||
&& warn "Stranded rule removed — remote connectivity restored ✅" \
|
||||
|| error "Could not remove stranded rule — run: iptables -D OUTPUT -d $REMOTE_SERVER -j DROP"
|
||||
fi
|
||||
exit 0
|
||||
fi
|
||||
warn "Stopping fallback_test.sh (PID $target_pid) — SIGTERM so its trap clears the iptables rule..."
|
||||
kill -TERM "$target_pid" 2>/dev/null || true
|
||||
waited=0
|
||||
while kill -0 "$target_pid" 2>/dev/null && [[ "$waited" -lt 30 ]]; do
|
||||
sleep 1
|
||||
(( waited++ )) || true
|
||||
done
|
||||
if kill -0 "$target_pid" 2>/dev/null; then
|
||||
error "fallback_test.sh (PID $target_pid) did not exit within 30s"
|
||||
error "NOT force-killing — SIGKILL would strand the iptables DROP rule on $REMOTE_SERVER"
|
||||
error "Wait, or clear manually: iptables -D OUTPUT -d $REMOTE_SERVER -j DROP"
|
||||
exit 1
|
||||
fi
|
||||
warn "Stopped: fallback_test.sh (PID $target_pid) ✅"
|
||||
if iptables -C OUTPUT -d "${REMOTE_SERVER:-0.0.0.0}" -j DROP 2>/dev/null; then
|
||||
error "iptables DROP rule for $REMOTE_SERVER survived the stop — removing"
|
||||
iptables -D OUTPUT -d "$REMOTE_SERVER" -j DROP 2>/dev/null \
|
||||
&& warn "Rule removed ✅" || error "Could not remove — run it by hand"
|
||||
else
|
||||
log "No iptables DROP rule remains for $REMOTE_SERVER ✅"
|
||||
fi
|
||||
exit 0
|
||||
fi
|
||||
|
||||
FALLBACK_SCRIPT="$SCRIPT_DIR/fallback.sh"
|
||||
DOCKER_TIMEOUT=15
|
||||
|
||||
@@ -139,6 +210,28 @@ cleanup() {
|
||||
|
||||
trap cleanup EXIT
|
||||
|
||||
# ── Stale rule sweep — the one case the trap above cannot cover ───────────────────────────────
|
||||
# The EXIT trap fires on normal exit, error, ctrl-c and SIGTERM, but not on SIGKILL or a power
|
||||
# cut. A DROP rule stranded that way makes fallback.sh see the partner as permanently down and
|
||||
# sit in FALLBACK indefinitely — so clear any leftover from a previous run before starting.
|
||||
_clear_stale_block() {
|
||||
[[ -z "${REMOTE_SERVER:-}" ]] && return
|
||||
local removed=0
|
||||
while iptables -C OUTPUT -d "$REMOTE_SERVER" -j DROP 2>/dev/null; do
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would remove stale iptables block on $REMOTE_SERVER"
|
||||
return
|
||||
fi
|
||||
iptables -D OUTPUT -d "$REMOTE_SERVER" -j DROP 2>/dev/null || break
|
||||
(( removed++ ))
|
||||
done
|
||||
if [[ "$removed" -gt 0 ]]; then
|
||||
warn "$ICON_SHIELD Removed $removed stale iptables block(s) on $REMOTE_SERVER from a previous run"
|
||||
notify "Fallback test cleared $removed stale iptables block(s) on $(hostname) — a previous test was killed before cleanup" \
|
||||
"Fallback Test" "warning"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
@@ -148,8 +241,9 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# FALLBACK_ENABLED gate — no point testing if failover is disabled
|
||||
if [[ "${FALLBACK_ENABLED:-false}" == false ]]; then
|
||||
# FALLBACK_ENABLED gate — no point testing if fallback is disabled
|
||||
# Fail-closed, matching fallback.sh — anything not exactly "true" counts as disabled.
|
||||
if [[ "${FALLBACK_ENABLED:-false}" != "true" ]]; then
|
||||
warn "FALLBACK_ENABLED=false — fallback test aborted"
|
||||
warn "Enable fallback in master.conf before running this test"
|
||||
exit 0
|
||||
@@ -166,21 +260,18 @@ detect_hosts
|
||||
resolve_remote_ip
|
||||
|
||||
# Validate commands used by this script
|
||||
validate_unraid_cmd \
|
||||
platform_require_cmd \
|
||||
"$(which iptables 2>/dev/null || echo /sbin/iptables)" \
|
||||
"--version" "iptables" \
|
||||
"iptables" || { error "iptables not found — required for connectivity simulation"; exit 1; }
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
if [[ ! -f "$FALLBACK_SCRIPT" ]]; then
|
||||
error "fallback.sh not found at $FALLBACK_SCRIPT"
|
||||
exit 1
|
||||
fi
|
||||
log "fallback.sh found at $FALLBACK_SCRIPT"
|
||||
log "$ICON_GEAR Config: remote=${REMOTE_SERVER_NAME} (${REMOTE_SERVER}) fallback-script=${FALLBACK_SCRIPT}"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no iptables rules or container changes will be made"
|
||||
|
||||
@@ -188,13 +279,13 @@ log "fallback.sh found at $FALLBACK_SCRIPT"
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
local_ver=$(grep -oP '(?<=version=")[^"]+' /etc/unraid-version 2>/dev/null || echo "unknown")
|
||||
local_ver=$(platform_get_os_version 2>/dev/null || echo "unknown")
|
||||
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST My ID: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_HOST Remote ID: $REMOTE_ID ($REMOTE_SERVER_NAME — $REMOTE_SERVER)"
|
||||
echo "$ICON_GEAR unRAID ver: $local_ver"
|
||||
echo "$ICON_GEAR OS ver: $local_ver"
|
||||
echo "$ICON_FALLBACK Block wait: ${FALLBACK_TEST_BLOCK_WAIT}s"
|
||||
echo "$ICON_FALLBACK Handback wait: ${FALLBACK_TEST_HANDBACK_WAIT}s"
|
||||
echo "$ICON_FALLBACK Check interval: ${FALLBACK_CHECK_INTERVAL}s"
|
||||
@@ -209,7 +300,7 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
fi
|
||||
|
||||
# Show Tier 1 containers for this host
|
||||
TIER1_VAR="FALLBACK_${MY_ID}_COVERS_${REMOTE_ID}_TIER1"
|
||||
TIER1_VAR="FALLBACK_${REMOTE_ID}_TIER1"
|
||||
eval "TIER1_CONTAINERS=(\"\${${TIER1_VAR}[@]:-}\")"
|
||||
echo "$ICON_CONTAINERS Tier 1 to test: ${TIER1_CONTAINERS[*]:-none configured}"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
@@ -226,8 +317,8 @@ TOTAL_START=$(date +%s)
|
||||
phase_pass() { PHASES_PASS+=("$1"); warn "$ICON_DONE Phase: $1 — PASSED ✅"; }
|
||||
phase_fail() { PHASES_FAIL+=("$1"); error "Phase: $1 — FAILED ❌"; }
|
||||
|
||||
# Get Tier 1 containers for this server's failover responsibility
|
||||
TIER1_VAR="FALLBACK_${MY_ID}_COVERS_${REMOTE_ID}_TIER1"
|
||||
# Get Tier 1 containers defined in remote's own conf
|
||||
TIER1_VAR="FALLBACK_${REMOTE_ID}_TIER1"
|
||||
eval "TIER1_CONTAINERS=(\"\${${TIER1_VAR}[@]:-}\")"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -241,6 +332,11 @@ echo "━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Phase 1 — Pre-flight ━━━"
|
||||
|
||||
# Must run before the reachability check below — a stale DROP rule from a killed run makes
|
||||
# the partner look unreachable, and the test would abort blaming the network for its own
|
||||
# leftover.
|
||||
_clear_stale_block
|
||||
|
||||
# Remote reachable
|
||||
if ping_remote; then
|
||||
log "$REMOTE_SERVER_NAME is reachable"
|
||||
@@ -260,7 +356,7 @@ else
|
||||
fi
|
||||
|
||||
# Version parity — test may produce misleading results on mismatch
|
||||
if ! check_unraid_version_parity; then
|
||||
if ! check_os_version_parity; then
|
||||
error "unRAID version mismatch — test aborted to prevent misleading results"
|
||||
phase_fail "Pre-flight"
|
||||
exit 1
|
||||
@@ -286,10 +382,31 @@ else
|
||||
warn "No state file found — assuming NORMAL (first run)"
|
||||
fi
|
||||
|
||||
# fallback.sh must actually be RUNNING, not merely enabled
|
||||
#
|
||||
# Every phase after this one waits for the daemon to change state. FALLBACK_ENABLED=true says
|
||||
# it is allowed to run; it does not say array_started.sh launched it, or that it is still alive.
|
||||
# Without this the test passes pre-flight, drops a real iptables rule on the partner, waits
|
||||
# FALLBACK_TEST_BLOCK_WAIT for a transition nothing is there to make, and fails Phase 3 blaming
|
||||
# fallback detection. Only the EXIT trap gets connectivity back.
|
||||
#
|
||||
# In --dry-run nothing is blocked and nothing is waited on, so a dead daemon is worth saying but
|
||||
# not worth aborting for — the walkthrough still shows the operator the shape of the run.
|
||||
if pgrep -f "Fallback/fallback\.sh" >/dev/null 2>&1; then
|
||||
log "fallback.sh daemon is running"
|
||||
elif [[ "$DRY_RUN" == true ]]; then
|
||||
warn "fallback.sh is NOT running — a real test would abort here"
|
||||
else
|
||||
error "fallback.sh is not running — nothing would detect the outage this test creates"
|
||||
error "Start it with array_started.sh, or run with --dry-run to walk the phases"
|
||||
phase_fail "Pre-flight"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Tier 1 containers configured
|
||||
if [[ ${#TIER1_CONTAINERS[@]} -eq 0 ]]; then
|
||||
error "No Tier 1 containers configured for $MY_ID → $REMOTE_ID"
|
||||
error "Check FALLBACK_${MY_ID}_COVERS_${REMOTE_ID}_TIER1 in host*.conf"
|
||||
error "Check FALLBACK_${REMOTE_ID}_TIER1 in ${REMOTE_ID,,}.conf"
|
||||
phase_fail "Pre-flight"
|
||||
exit 1
|
||||
fi
|
||||
@@ -337,7 +454,7 @@ if [[ "$DRY_RUN" == false ]]; then
|
||||
if [[ -f "$FALLBACK_STATE_FILE" ]]; then
|
||||
NEW_STATE=$(grep "^state=" "$FALLBACK_STATE_FILE" 2>/dev/null | cut -d= -f2)
|
||||
if [[ "$NEW_STATE" == "FALLBACK" ]]; then
|
||||
log "State changed to FALLBACK — outage detected correctly ✅"
|
||||
echo "State changed to FALLBACK — outage detected correctly ✅"
|
||||
phase_pass "Fallback Detection"
|
||||
else
|
||||
error "State is $NEW_STATE — expected FALLBACK after ${FALLBACK_TEST_BLOCK_WAIT}s"
|
||||
@@ -367,7 +484,7 @@ if [[ "$DRY_RUN" == false ]]; then
|
||||
STATUS=$(timeout "$DOCKER_TIMEOUT" docker inspect -f '{{.State.Running}}' \
|
||||
"$container" 2>/dev/null)
|
||||
if [[ "$STATUS" == "true" ]]; then
|
||||
log "$ICON_RUNNING $container is running locally ✅"
|
||||
echo "$ICON_RUNNING $container is running locally ✅"
|
||||
else
|
||||
error "$ICON_NOT_RUNNING $container is NOT running locally"
|
||||
CONTAINERS_OK=false
|
||||
@@ -397,7 +514,7 @@ if [[ "$DRY_RUN" == false ]]; then
|
||||
|
||||
sleep 3
|
||||
if ping_remote; then
|
||||
log "$REMOTE_SERVER_NAME is reachable again ✅"
|
||||
echo "$REMOTE_SERVER_NAME is reachable again ✅"
|
||||
phase_pass "Restore Connectivity"
|
||||
else
|
||||
error "$REMOTE_SERVER_NAME still unreachable after removing iptables rule"
|
||||
@@ -423,7 +540,7 @@ if [[ "$DRY_RUN" == false ]]; then
|
||||
if [[ -f "$FALLBACK_STATE_FILE" ]]; then
|
||||
FINAL_STATE=$(grep "^state=" "$FALLBACK_STATE_FILE" 2>/dev/null | cut -d= -f2)
|
||||
if [[ "$FINAL_STATE" == "NORMAL" ]]; then
|
||||
log "State returned to NORMAL — handback completed ✅"
|
||||
echo "State returned to NORMAL — handback completed ✅"
|
||||
phase_pass "Handback"
|
||||
else
|
||||
error "State is $FINAL_STATE — expected NORMAL after ${FALLBACK_TEST_HANDBACK_WAIT}s"
|
||||
@@ -453,7 +570,7 @@ if [[ "$DRY_RUN" == false ]]; then
|
||||
STATUS=$(timeout "$DOCKER_TIMEOUT" docker inspect -f '{{.State.Running}}' \
|
||||
"$container" 2>/dev/null)
|
||||
if [[ "$STATUS" != "true" ]]; then
|
||||
log "$ICON_NOT_RUNNING $container stopped locally — handed back ✅"
|
||||
echo "$ICON_NOT_RUNNING $container stopped locally — handed back ✅"
|
||||
else
|
||||
error "$ICON_RUNNING $container still running locally — handback may have failed"
|
||||
HANDBACK_OK=false
|
||||
|
||||
Regular → Executable
+76
-2
@@ -12,7 +12,21 @@
|
||||
# Does NOT download media, search indexers, or manage applications. Has no
|
||||
# side effects — no file writes, no API calls, no deletions.
|
||||
#
|
||||
# Currently paired with: Media/playback_aware_lidarr_discovery.sh
|
||||
# Currently paired with: Arrs_Stack/playback_aware_lidarr_discovery.sh
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# A library, not a program. Consumers source it and drive the pipeline themselves:
|
||||
#
|
||||
# 1. score_candidate() — sum the four weighted component scores
|
||||
# 2. apply_temporal_decay() — reduce by one unit per 30 days of age, floored at 0
|
||||
# 3. is_duplicate_candidate() — check the consumer's own history file
|
||||
# 4. make_decision() — ACCEPT or REJECT against the consumer's threshold
|
||||
#
|
||||
# Every input is supplied by the caller and every output is returned to it. The engine holds
|
||||
# no state between calls, reads no config, and never acts on its own verdict.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
@@ -38,6 +52,60 @@
|
||||
# the consumer's concern — the engine never acts on its own verdict.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# No Side Effects — Structural, Not Incidental
|
||||
# Writes no files, makes no API calls, deletes nothing, starts nothing. This is the
|
||||
# safeguard: a scoring mistake here can only ever produce a wrong number, never a wrong
|
||||
# action. Keep it that way — the moment this library acquires a side effect, every consumer
|
||||
# inherits it silently.
|
||||
#
|
||||
# No Root, No Lock, No detect_hosts — Deliberate
|
||||
# Correct for a sourced library and should not be "fixed" to match the executable scripts.
|
||||
# There is no state for a lock to protect, no privileged operation to gate, and no
|
||||
# host-specific config to alias. It runs entirely inside the caller's process.
|
||||
#
|
||||
# Caller Owns the Verdict
|
||||
# ACCEPT/REJECT is a return value, not an instruction. Nothing here can cause a candidate
|
||||
# to be added, removed, or downloaded — the consumer decides what a verdict means.
|
||||
#
|
||||
# Literal Duplicate Matching
|
||||
# is_duplicate_candidate() matches with grep -Fx, treating the candidate as data rather
|
||||
# than a pattern. Artist and title strings routinely contain regex metacharacters, and a
|
||||
# false positive here silently discards a genuinely new candidate.
|
||||
#
|
||||
# Decay Floors at Zero
|
||||
# apply_temporal_decay() clamps at 0, so an old signal can never become a negative score
|
||||
# that drags an otherwise-passing candidate below threshold.
|
||||
#
|
||||
# Integer Arithmetic Throughout
|
||||
# All scoring is integer. No floating point means no locale-dependent decimal parsing and
|
||||
# no rounding drift between hosts.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# None, by design. The engine reads no conf file and no environment variable.
|
||||
#
|
||||
# Thresholds, weights and strictness profiles live with the consumer — see
|
||||
# LIDARR_DISCOVERY_* / SONARR_DISCOVERY_* / RADARR_DISCOVERY_* in master.conf. That is what
|
||||
# keeps the core domain-agnostic: music strictness cannot leak into TV intake, because the
|
||||
# engine never learns which domain it is scoring for.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# None — this file is sourced, never executed:
|
||||
#
|
||||
# source "$SCRIPT_DIR/../Kernel/decision_engine.sh"
|
||||
#
|
||||
# It has no argument parsing, no --dry-run and no --status, because it takes no action that
|
||||
# a dry run could suppress.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# FUNCTIONS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -65,6 +133,7 @@ score_candidate() {
|
||||
local popularity_score="${2:-0}"
|
||||
local recency_score="${3:-0}"
|
||||
local quality_score="${4:-0}"
|
||||
local TOTAL_SCORE
|
||||
|
||||
TOTAL_SCORE=$(( \
|
||||
user_score + \
|
||||
@@ -110,7 +179,12 @@ is_duplicate_candidate() {
|
||||
local candidate="$1"
|
||||
local history_file="$2"
|
||||
|
||||
grep -qi "^${candidate}$" "$history_file" 2>/dev/null
|
||||
# -F -x, not "^$" anchors: the candidate is data, not a pattern. Interpolating it into a
|
||||
# regex makes every metacharacter in an artist or title active — "R.E.M." matches "RxExMy",
|
||||
# and an unbalanced bracket makes grep error out entirely. Either way the caller reads the
|
||||
# result as "already seen" and silently skips something genuinely new. -F disables regex,
|
||||
# -x anchors the whole line, which is exactly what the anchors were reaching for.
|
||||
grep -qiFx -- "$candidate" "$history_file" 2>/dev/null
|
||||
}
|
||||
|
||||
# ── make_decision ─────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -3,11 +3,23 @@
|
||||
Getting a fresh two-server ecosystem running from scratch.
|
||||
For system overview see [README.md](README.md). For individual subsystem detail see folder READMEs and script headers.
|
||||
|
||||
**Ten steps, and they're in this order for a reason.** Each one assumes the previous one
|
||||
actually worked — not that you ran it, that it *worked*. Every step below ends with a way to
|
||||
check, and skipping those checks is how you end up three steps later debugging the wrong
|
||||
thing entirely.
|
||||
|
||||
The steps that hurt most when rushed are **3** (naming) and **8** (testing failover). Step 3
|
||||
because renaming anything afterwards means chasing it through every conf, and step 8 because
|
||||
untested failover isn't redundancy — it's a guess you haven't checked yet.
|
||||
|
||||
Budget an evening. It is not hard, but it is not five minutes either.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ BEFORE YOU START ━━━
|
||||
|
||||
Things that must be in place before you touch any scripts.
|
||||
None of this is optional and none of it is Varaverk's job to install. Get these in place
|
||||
first — every step after here assumes they're already true.
|
||||
|
||||
---
|
||||
|
||||
@@ -30,7 +42,6 @@ Install these via **Apps** (Community Applications) on both servers:
|
||||
| Plugin | Why |
|
||||
|--------|-----|
|
||||
| **Tailscale** | Encrypted VPN between servers — all script traffic travels over it |
|
||||
| **User Scripts** | Schedules the orchestrators (replaces per-script cron entries) |
|
||||
|
||||
---
|
||||
|
||||
@@ -89,8 +100,8 @@ hostname
|
||||
### ── Clone on HOST1 ───────────────────────────────────────────────────────────
|
||||
|
||||
```bash
|
||||
git clone <your_repo_url> /mnt/user/appdata/unraid_scripts
|
||||
cd /mnt/user/appdata/unraid_scripts
|
||||
git clone <your_repo_url> /boot/config/plugins/varaverk
|
||||
cd /boot/config/plugins/varaverk
|
||||
```
|
||||
|
||||
---
|
||||
@@ -98,8 +109,8 @@ cd /mnt/user/appdata/unraid_scripts
|
||||
### ── Clone on HOST2 ───────────────────────────────────────────────────────────
|
||||
|
||||
```bash
|
||||
git clone <your_repo_url> /mnt/user/appdata/unraid_scripts
|
||||
cd /mnt/user/appdata/unraid_scripts
|
||||
git clone <your_repo_url> /boot/config/plugins/varaverk
|
||||
cd /boot/config/plugins/varaverk
|
||||
```
|
||||
|
||||
---
|
||||
@@ -112,7 +123,7 @@ credentials. Set up sparse checkout so each server only receives its own conf fi
|
||||
|
||||
**On HOST1** — exclude host2.conf:
|
||||
```bash
|
||||
cd /mnt/user/appdata/unraid_scripts
|
||||
cd /boot/config/plugins/varaverk
|
||||
git sparse-checkout init --no-cone
|
||||
git sparse-checkout set '/*' '!/Configurations/host2.conf'
|
||||
git checkout
|
||||
@@ -120,7 +131,7 @@ git checkout
|
||||
|
||||
**On HOST2** — exclude host1.conf:
|
||||
```bash
|
||||
cd /mnt/user/appdata/unraid_scripts
|
||||
cd /boot/config/plugins/varaverk
|
||||
git sparse-checkout init --no-cone
|
||||
git sparse-checkout set '/*' '!/Configurations/host1.conf'
|
||||
git checkout
|
||||
@@ -158,6 +169,21 @@ Replace these with your actual unRAID server hostnames. These must match exactly
|
||||
`detect_hosts()` compares the running server's hostname against these two values to
|
||||
know which server it is on. Everything else in the ecosystem flows from this.
|
||||
|
||||
**Get this wrong and nothing works, but nothing errors either.** If neither value matches,
|
||||
`MY_ID` is never set, and every script that depends on it either exits early or resolves
|
||||
`${MY_ID}_SOMETHING` to an empty variable and quietly takes the wrong branch. Copy the value
|
||||
straight out of `hostname` on each box rather than typing what you think it is:
|
||||
|
||||
```bash
|
||||
hostname # run this on each server, paste the exact output
|
||||
```
|
||||
|
||||
> **The 15-character trap.** unRAID truncates the Server Name to 15 characters for NetBIOS.
|
||||
> If your name is longer, what you set in the WebGUI and what `hostname` returns are two
|
||||
> different strings. `detect_hosts()` has a fallback that matches a truncated 15-char
|
||||
> hostname against a longer configured value — but only when the live hostname is
|
||||
> *exactly* 15 characters. Don't rely on it. Use the real `hostname` output.
|
||||
|
||||
---
|
||||
|
||||
### ── 3b. host1.conf — HOST1 identity ────────────────────────────────────────
|
||||
@@ -207,15 +233,27 @@ HOST1_CRITICAL_SYNC_SHARES=(
|
||||
"/mnt/user/appdata-Fallback/Critical-Data|critical-data"
|
||||
)
|
||||
|
||||
# Containers fallback.sh starts on HOST2 when HOST1 goes down (Tier 1 = immediate)
|
||||
HOST2_FALLBACK_HOST1_COVERS_HOST2_TIER1=() # HOST2's own covers for HOST1, and vice versa
|
||||
HOST1_FALLBACK_HOST2_COVERS_HOST1_TIER1=(
|
||||
# Containers to start when HOST1 goes down (Tier 1 = immediate)
|
||||
# Lives in host1.conf. Named for the host being COVERED, not the host doing the covering.
|
||||
FALLBACK_HOST1_TIER1=(
|
||||
"Emby"
|
||||
"VaultWarden"
|
||||
"NginxProxyManager"
|
||||
)
|
||||
```
|
||||
|
||||
**Read that naming carefully — it catches people.** The variable is
|
||||
`FALLBACK_${REMOTE_ID}_TIER${N}`, and `REMOTE_ID` is *the host that went down*. So
|
||||
`FALLBACK_HOST1_TIER1` is the list HOST2 reads when HOST1 is unreachable. It defines
|
||||
"what HOST1 needs covering", not "what HOST2 runs".
|
||||
|
||||
It lives in `host1.conf` for the same reason — HOST1 owns the description of its own
|
||||
service stack. HOST2 receives it through the conf cache rather than keeping its own
|
||||
opinion about what HOST1 runs. One list, one owner, no drift.
|
||||
|
||||
Tiers 2–4 activate progressively after `HOST1_TIER2_DELAY` etc., so a brief outage never
|
||||
drags the entire remote stack across.
|
||||
|
||||
For full configuration reference see `Manual-Fallback.md`, `Manual-Rsync.md`,
|
||||
and `Manual-Watchdogs.md`.
|
||||
|
||||
@@ -235,7 +273,7 @@ prompted for HOST1's root password once (for the initial key copy). After that,
|
||||
all SSH is keyless.
|
||||
|
||||
```bash
|
||||
cd /mnt/user/appdata/unraid_scripts
|
||||
cd /boot/config/plugins/varaverk
|
||||
bash Partnership/ssh_setup.sh
|
||||
```
|
||||
|
||||
@@ -249,7 +287,7 @@ When prompted, enter HOST1's root password. The script:
|
||||
### ── Run on HOST1 ──────────────────────────────────────────────────────────
|
||||
|
||||
```bash
|
||||
cd /mnt/user/appdata/unraid_scripts
|
||||
cd /boot/config/plugins/varaverk
|
||||
bash Partnership/ssh_setup.sh
|
||||
```
|
||||
|
||||
@@ -296,49 +334,59 @@ ssh -o StrictHostKeyChecking=accept-new root@unRAID-Gmer4Lfe "echo ok"
|
||||
|
||||
---
|
||||
|
||||
## ━━━ STEP 5: ARBITER — USER SCRIPTS SETUP ━━━
|
||||
## ━━━ STEP 5: SCHEDULER SETUP ━━━
|
||||
|
||||
Arbiter (the User Scripts plugin) is the only thing Arbiter runs — never individual
|
||||
scripts directly. Each entry is an orchestrator that calls everything in its window.
|
||||
**Everything runs through the Varaverk plugin's built-in scheduler** — no User Scripts
|
||||
entries needed. The plugin handles all triggers natively:
|
||||
|
||||
Open **Settings → User Scripts** in the unRAID WebGUI.
|
||||
- **Array start** → `Plugin/unraid/event/disks_mounted/array_start_jobs` fires `array_started.sh`
|
||||
- **Array stop** → `Plugin/unraid/event/disks_unmounting/array_stop_jobs` fires `array_stopping.sh`
|
||||
- **Cron** → `Plugin/unraid/event/disks_mounted/rebuild_cron` rebuilds the cron file from `schedule.json` on every boot
|
||||
|
||||
Configure via the Varaverk plugin Scheduler tab (or edit `schedule.json` directly).
|
||||
The built-in job list schedules orchestrators, never individual repo scripts directly —
|
||||
but the Scheduler tab's **Custom Scripts** card is the one place individual scripts
|
||||
*are* scheduled directly (see below).
|
||||
|
||||
---
|
||||
|
||||
### ── Required entries (both servers) ────────────────────────────────────────
|
||||
### ── Custom Scripts ────────────────────────────────────────────────────────────
|
||||
|
||||
Create one User Script entry for each row. Set the schedule in the cron field and
|
||||
paste the command. Name the entry to match the script name.
|
||||
The Scheduler tab has a **Custom Scripts** card for one-off scripts that aren't part of
|
||||
the repo's orchestrator pipeline — personal tooling, quick fixes, anything you don't
|
||||
want to wire into `master.conf`.
|
||||
|
||||
| Script | Cron schedule | Command |
|
||||
|--------|--------------|---------|
|
||||
| `array_started` | At Array Start | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/array_started.sh` |
|
||||
| `watchdog_orchestrator` | `* * * * *` | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/watchdog_orchestrator.sh` |
|
||||
| `transcode_management` | `*/7 * * * *` | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/transcode_management.sh` |
|
||||
| `critical_sync_maintenance` | `*/30 * * * *` | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/critical_sync_maintenance.sh` |
|
||||
| `emby_fallback_sync` | `*/30 * * * *` | `bash /mnt/user/appdata/unraid_scripts/Rsync/rsync.sh /mnt/user/Media_Server/Emby --profile=emby-fallback` |
|
||||
| `intermediate_sync_maintenance` | `0 */4 * * *` | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/intermediate_sync_maintenance.sh` |
|
||||
| `daily_sync_maintenance` | `0 1 * * *` | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/daily_sync_maintenance.sh` |
|
||||
| `weekly_sync_maintenance` | `30 2 * * 0` | `bash /mnt/user/appdata/unraid_scripts/Orchestrators/weekly_sync_maintenance.sh` |
|
||||
| `weekly_health_digest` | `0 8 * * *` | `bash /mnt/user/appdata/unraid_scripts/Monitors/weekly_health_digest.sh` |
|
||||
Scripts live in `/boot/config/plugins/user.scripts/Varaverk/Scripts/` — deliberately
|
||||
**outside** the Varaverk git repo (that folder is never pushed to GitHub), in the same
|
||||
place the Unraid User Scripts plugin keeps its own scripts, so it's a folder location
|
||||
admins are already used to.
|
||||
|
||||
> `array_started` uses the built-in "At Array Start" schedule option in User Scripts —
|
||||
> it is not a cron expression.
|
||||
Two ways to get a script there:
|
||||
|
||||
> `emby_fallback_sync` is the only script with a direct Arbiter entry. It is separate
|
||||
> from `critical_sync_maintenance` because it needs its own 30-minute slot and Emby
|
||||
> must stay running during this sync (dirty sync — no container stop).
|
||||
- Click **+ Create Script** on the Scheduler tab — opens an inline editor, writes the
|
||||
file to that folder, and adds a `schedule.json` entry automatically.
|
||||
- Drop any `.sh` file into the folder yourself (e.g. via terminal, or Unraid's own
|
||||
Custom Scripts / User Scripts plugin pointed at the same path). The Scheduler tab
|
||||
**auto-detects** it — discovery is a folder scan, not a registry, so it doesn't matter
|
||||
how the file got there. It shows up disabled with no cron until you configure one.
|
||||
|
||||
---
|
||||
|
||||
### ── Schedule note ────────────────────────────────────────────────────────────
|
||||
### ── Schedule (Varaverk Scheduler) ───────────────────────────────────────────
|
||||
|
||||
`transcode_management` and `critical_sync_maintenance` both run every 7 and 30 minutes
|
||||
respectively. In User Scripts, concurrent runs of the same script are prevented by the
|
||||
lock system (`acquire_lock`) — if the previous run is still active the new one exits
|
||||
immediately. No need to stagger them manually.
|
||||
| Script | Event/Cron | Purpose |
|
||||
|--------|-----------|---------|
|
||||
| `array_started` | `array_start` | All array startup scripts in order |
|
||||
| `array_stopping` | `array_stop` | Ordered graceful shutdown |
|
||||
| `transcode_management` | `*/7 * * * *` | Cleanup then manager — order critical |
|
||||
| `watchdog_orchestrator` | `*/15 * * * *` | resource → docker → system → stability |
|
||||
| `critical_sync_maintenance` | `*/30 * * * *` | auth + Emby dirty sync + partnership |
|
||||
| `intermediate_sync_maintenance` | `0 */4 * * *` | arr sync + failed recovery |
|
||||
| `daily_sync_maintenance` | `0 1 * * *` | full daily maintenance window |
|
||||
| `weekly_sync_maintenance` | `30 2 * * 0` | clean sync + image updates |
|
||||
| `monthly_maintenance` | `0 0 15 * *` | ZFS scrub, SMART tests (uptime-gated) |
|
||||
|
||||
For the complete schedule and what each orchestrator runs, see
|
||||
Concurrent runs are prevented by `acquire_lock`. For the complete schedule see
|
||||
[README-Orchestrators.md](Orchestrators/README-Orchestrators.md).
|
||||
|
||||
---
|
||||
@@ -354,7 +402,7 @@ will take a while depending on library size.
|
||||
|
||||
```bash
|
||||
# On HOST1 — preview what would sync, check paths and profiles
|
||||
bash /mnt/user/appdata/unraid_scripts/Orchestrators/daily_sync_maintenance.sh --dry-run --log
|
||||
bash /boot/config/plugins/varaverk/Orchestrators/daily_sync_maintenance.sh --dry-run --log
|
||||
```
|
||||
|
||||
Review the output. Check that:
|
||||
@@ -368,11 +416,11 @@ Review the output. Check that:
|
||||
|
||||
```bash
|
||||
# On HOST1 — run the full daily sync (this includes rsync for all DAILY_SYNC_SHARES)
|
||||
bash /mnt/user/appdata/unraid_scripts/Orchestrators/daily_sync_maintenance.sh --log
|
||||
bash /boot/config/plugins/varaverk/Orchestrators/daily_sync_maintenance.sh --log
|
||||
```
|
||||
|
||||
This will take longer than future nightly runs — it is transferring everything for
|
||||
the first time. Monitor progress via the User Scripts output or:
|
||||
the first time. Monitor progress via the Varaverk log viewer or:
|
||||
|
||||
```bash
|
||||
# Watch transfer in real time
|
||||
@@ -381,7 +429,7 @@ watch -n5 'ls -lh /mnt/user/Movies/ | tail -5'
|
||||
|
||||
After it finishes, also run the critical sync to populate auth and Emby state:
|
||||
```bash
|
||||
bash /mnt/user/appdata/unraid_scripts/Orchestrators/critical_sync_maintenance.sh --log
|
||||
bash /boot/config/plugins/varaverk/Orchestrators/critical_sync_maintenance.sh --log
|
||||
```
|
||||
|
||||
---
|
||||
@@ -410,7 +458,7 @@ Commit and pull on both servers so both pick up the change.
|
||||
|
||||
To start it now without a reboot:
|
||||
```bash
|
||||
bash /mnt/user/appdata/unraid_scripts/Fallback/fallback.sh &
|
||||
bash /boot/config/plugins/varaverk/Fallback/fallback.sh &
|
||||
```
|
||||
|
||||
---
|
||||
@@ -424,28 +472,40 @@ pgrep -a -f fallback.sh
|
||||
|
||||
Check current state:
|
||||
```bash
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
# state=NORMAL — both servers up
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ STEP 8: TEST FAILOVER ━━━
|
||||
## ━━━ STEP 8: TEST FALLBACK ━━━
|
||||
|
||||
Before relying on the system, confirm it actually triggers. `fallback_test.sh`
|
||||
simulates an outage using `iptables` — no real downtime, no real data changes.
|
||||
|
||||
```bash
|
||||
# On HOST2 — dry run first (sees the sequence, no container starts)
|
||||
bash /mnt/user/appdata/unraid_scripts/Fallback/fallback_test.sh --dry-run --log
|
||||
bash /boot/config/plugins/varaverk/Fallback/fallback_test.sh --dry-run --log
|
||||
|
||||
# When ready — real test (uses iptables to simulate HOST1 unreachable)
|
||||
bash /mnt/user/appdata/unraid_scripts/Fallback/fallback_test.sh --log
|
||||
bash /boot/config/plugins/varaverk/Fallback/fallback_test.sh --log
|
||||
```
|
||||
|
||||
The test runs in phases — blocks HOST1's Tailscale IP, waits for `fallback.sh` to
|
||||
detect it and start Tier 1 containers, then unblocks and waits for handback.
|
||||
|
||||
**Do not skip this step.** Everything up to here you can verify by reading output. Failover
|
||||
is the one part you cannot confirm by looking at it — a typo in a tier list, a container name
|
||||
that doesn't exist on the other side, a DDNS container that was renamed six months ago: all
|
||||
of it sits there looking completely fine until the moment it's needed.
|
||||
|
||||
The test costs you twenty minutes and no downtime. The alternative is finding out at 2am,
|
||||
during the outage, when you have exactly one chance to get it right.
|
||||
|
||||
**Watch the handback as closely as the failover.** Coming back is the harder half — it has to
|
||||
stop the remote's DDNS, sync the data back, start the local containers, and only then bring
|
||||
local DDNS up. Failover starting correctly tells you nothing about whether handback does.
|
||||
|
||||
See [Manual-Fallback.md](Fallback/Manual-Fallback.md) for what each phase does and
|
||||
how to interpret the output.
|
||||
|
||||
@@ -478,9 +538,9 @@ HOST2 runs first — it generates its SSH key and waits. The owner completes set
|
||||
|
||||
```bash
|
||||
# On HOST2
|
||||
bash /mnt/user/appdata/unraid_scripts/Partnership/partnership_onboard.sh --dry-run --log
|
||||
bash /boot/config/plugins/varaverk/Partnership/partnership_onboard.sh --dry-run --log
|
||||
# Review output, then:
|
||||
bash /mnt/user/appdata/unraid_scripts/Partnership/partnership_onboard.sh --log
|
||||
bash /boot/config/plugins/varaverk/Partnership/partnership_onboard.sh --log
|
||||
```
|
||||
|
||||
HOST2's work is done after Step 1 (SSH key setup). The rest happens on HOST1.
|
||||
@@ -491,9 +551,9 @@ HOST2's work is done after Step 1 (SSH key setup). The rest happens on HOST1.
|
||||
|
||||
```bash
|
||||
# On HOST1
|
||||
bash /mnt/user/appdata/unraid_scripts/Partnership/partnership_onboard.sh --dry-run --log
|
||||
bash /boot/config/plugins/varaverk/Partnership/partnership_onboard.sh --dry-run --log
|
||||
# Review output — this will deploy auth + arr stacks to HOST2 remotely
|
||||
bash /mnt/user/appdata/unraid_scripts/Partnership/partnership_onboard.sh --log
|
||||
bash /boot/config/plugins/varaverk/Partnership/partnership_onboard.sh --log
|
||||
```
|
||||
|
||||
The owner path deploys containers to HOST2 via SSH, configures WebUI targets, runs
|
||||
@@ -525,7 +585,7 @@ bash Orchestrators/daily_sync_maintenance.sh --dry-run --log
|
||||
### ── Check fallback state ────────────────────────────────────────────────────
|
||||
|
||||
```bash
|
||||
cat /boot/config/fallback_state.db
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
# Expected:
|
||||
# state=NORMAL
|
||||
# fallback_start=0
|
||||
@@ -542,8 +602,8 @@ cat /boot/config/fallback_state.db
|
||||
Bandwidth monitor records every sync:
|
||||
```bash
|
||||
# Latest transfer log
|
||||
ls -lt /mnt/user/appdata/unraid_scripts/data/bandwidth_monitor/
|
||||
cat /mnt/user/appdata/unraid_scripts/data/bandwidth_monitor/<latest_log>
|
||||
ls -lt /boot/config/plugins/varaverk/data/bandwidth_monitor/
|
||||
cat /boot/config/plugins/varaverk/data/bandwidth_monitor/<latest_log>
|
||||
```
|
||||
|
||||
---
|
||||
@@ -615,8 +675,8 @@ The remote server is unreachable on the first check. Common causes:
|
||||
|
||||
```bash
|
||||
# Confirm which state fallback is in
|
||||
cat /boot/config/fallback_state.db
|
||||
# If stuck in FAILOVER after remote comes back: reset state
|
||||
cat "$STATE_DIR/fallback_state.db"
|
||||
# If stuck in FALLBACK after remote comes back: reset state
|
||||
bash Tools/fallback_state_reset.sh
|
||||
```
|
||||
|
||||
@@ -642,6 +702,14 @@ Each subsystem has a README with the design decisions and a Manual with the conf
|
||||
reference and troubleshooting. The `--status` flag on any script shows the current
|
||||
configuration and state.
|
||||
|
||||
**Turn things on one at a time and give each one a few days.** Everything below is off by
|
||||
default on a fresh install, and that's deliberate — a stack where six new subsystems went
|
||||
live the same night is a stack where you have no idea which one to blame. Enable, watch it
|
||||
through a full daily cycle, then enable the next.
|
||||
|
||||
Anything that deletes files — the arr cleanup scripts especially — gets a `--dry-run --log`
|
||||
first. Read the list. Every time, not just the first time.
|
||||
|
||||
```
|
||||
Watchdogs/ → README-Watchdogs.md configure memory limits, container lists
|
||||
Media/ → README-Media.md enable arr cleanup, discovery scripts
|
||||
|
||||
+68
-597
@@ -1,8 +1,11 @@
|
||||
# ━━━━━ MEDIA — Manual ━━━━━
|
||||
|
||||
Config reference, procedures, operational workflows.
|
||||
Config reference, procedures, operational workflows for Media/ scripts.
|
||||
For overview see README-Media.md. For per-script detail see script headers.
|
||||
|
||||
For arr stack config and procedures (cleanup, release fixer, sync, discovery) see
|
||||
`Arrs_Stack/Manual-Arrs_Stack.md`.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ PERMISSIONS MODEL ━━━
|
||||
@@ -32,48 +35,6 @@ Once fixed, this script corrects 0 files per run — it becomes a pure daily fai
|
||||
|
||||
---
|
||||
|
||||
## ━━━ ARR CLEANUP — FILE CLASSIFICATION ━━━
|
||||
|
||||
Every file found on disk during an arr cleanup run falls into exactly one category:
|
||||
|
||||
```
|
||||
TRACKED → arr API returned this exact path → leave it alone
|
||||
PROTECTED → matches ARR_PROTECTED_PATTERNS → never delete
|
||||
ORPHAN → media extension, not tracked, old enough → delete
|
||||
JUNK → not a media extension, not protected → delete (any age)
|
||||
RECENT → not tracked, under ARR_ORPHAN_AGE days → skip (may be mid-import)
|
||||
```
|
||||
|
||||
**Why protected patterns are critical:** arrs generate artwork (`*.jpg`), metadata
|
||||
(`*.nfo`), and subtitles/lyrics that do NOT appear in the tracked file API response.
|
||||
Without protection, these would be classified as orphans and deleted — removing cover art
|
||||
from every album, every movie poster, every TV show thumbnail. Requires a full rescan
|
||||
to recover. Never remove artwork extensions from protected patterns.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ ARR CLEANUP — SAFETY LAYERS ━━━
|
||||
|
||||
All 7 layers must pass before any file is touched. There is no way to push through a
|
||||
failed safety check without the explicit override flag.
|
||||
|
||||
```
|
||||
1. Container running + healthy — a stopped container has an empty API
|
||||
2. API reachable — no API = no tracked file list = everything looks orphaned
|
||||
3. API version matches — major version must match tested version in master.conf
|
||||
4. Item count > 0 — no artists/series/movies = something is wrong with DB
|
||||
5. Tracked file count > 0 — empty response = everything would be deleted
|
||||
6. Tracked count >= MIN_TRACKED_PCT — dramatic drop from last run = abort and alert
|
||||
7. Deletion size < MAX_DELETE_GB — last line of defense against misconfigured root path
|
||||
```
|
||||
|
||||
Layer 7 is the catastrophic failure prevention. A misconfigured root path — pointing
|
||||
cleanup at the wrong directory — means the API returns zero tracked files for a root
|
||||
that actually contains thousands. Everything walks as an orphan. Everything gets deleted.
|
||||
`LIDARR/SONARR/RADARR_MAX_DELETE_GB` requires `--i-know-what-im-doing` to proceed past it.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ CONFIGURATION — master.conf ━━━
|
||||
|
||||
### Permissions
|
||||
@@ -113,186 +74,51 @@ MEDIA_FILE_PATTERNS=(
|
||||
)
|
||||
```
|
||||
|
||||
> **DO NOT add patterns that match media you want to keep:**
|
||||
> `*.mkv *.mp4 *.avi *.m4v` — video files
|
||||
> `*.flac *.mp3 *.m4a` — audio files
|
||||
> `*.srt *.sub *.ass` — subtitle files (Bazarr managed)
|
||||
> `*.jpg *.png` — artwork
|
||||
> **DO NOT add patterns that match media you want to keep:**
|
||||
> `*.mkv *.mp4 *.avi *.m4v` — video files
|
||||
> `*.flac *.mp3 *.m4a` — audio files
|
||||
> `*.srt *.sub *.ass` — subtitle files (Bazarr managed)
|
||||
> `*.jpg *.png` — artwork
|
||||
> Always use `--dry-run` when adding new patterns.
|
||||
|
||||
---
|
||||
|
||||
### Lidarr Cleanup Thresholds
|
||||
### Play State Sync
|
||||
|
||||
```bash
|
||||
LIDARR_ORPHAN_AGE=7 # days — files newer than this are RECENT (mid-import window)
|
||||
LIDARR_MIN_TRACKED_PCT=80 # abort if API returns < 80% of last known count
|
||||
LIDARR_MAX_DELETE_GB=50 # require --i-know-what-im-doing above this
|
||||
LIDARR_IMPORT_SCAN_TIMEOUT=600 # seconds to wait for pre-flight import scan
|
||||
LIDARR_VERSION_MAJOR=3 # expected Lidarr major version (API safety check)
|
||||
LIDARR_EXTENSIONS=("flac" "mp3" "m4a" "wav" "aac" "ogg" "opus" "wma")
|
||||
LIDARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.lrc")
|
||||
LIDARR_TRACKED_COUNT_FILE=/boot/config/lidarr_tracked_count # persistent baseline
|
||||
ARR_CLEANUP_STATS=/boot/config/arr_cleanup_stats.db # read by coffee report
|
||||
PLAY_SYNC_ENABLED=true # toggle entire sync
|
||||
PLAY_SYNC_REMOTE=true # sync across hosts via Tailscale
|
||||
# false = local servers only (this host's Emby + Jellyfin)
|
||||
PLAY_SYNC_TYPES="Movie,Episode" # item types to sync — Audio excluded, music library too large
|
||||
PLAY_SYNC_FAV_TYPES="MusicArtist,MusicAlbum,Movie,Series" # favourites, union sync, never unmarks
|
||||
|
||||
PLAY_SYNC_PROBE=true # skip per-item work when nothing changed since last run
|
||||
PLAY_SYNC_PROBE_MAX_AGE_HOURS=24 # force a full comparison when the fingerprint is older
|
||||
|
||||
PLAY_SYNC_HANDBACK_RETRIES=5 # fallback handback: attempts before DNS cutover proceeds
|
||||
PLAY_SYNC_HANDBACK_RETRY_DELAY=60 # seconds between those attempts
|
||||
```
|
||||
|
||||
---
|
||||
> **There is no date/lookback filter, and one must not be re-added.** An earlier version
|
||||
> gated on played-date; it was removed once the real cost was measured — the 30-minute
|
||||
> runtime was fork overhead per item, not API volume or item count. The fix was jq epoch
|
||||
> parsing plus the response-hash probe below. Re-introducing a date window would reduce
|
||||
> correctness (older items silently stop syncing) without meaningfully reducing runtime.
|
||||
|
||||
### Sonarr Cleanup Thresholds
|
||||
Emby and Jellyfin servers configured per-host:
|
||||
|
||||
```bash
|
||||
SONARR_ORPHAN_AGE=7
|
||||
SONARR_MAX_DELETE_GB=50
|
||||
SONARR_IMPORT_SCAN_TIMEOUT=600
|
||||
SONARR_VERSION_MAJOR=4
|
||||
SONARR_EXTENSIONS=("mkv" "mp4" "avi" "m4v" "ts" "wmv" "mov")
|
||||
SONARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.ass" "*.ssa")
|
||||
```
|
||||
|
||||
Note: `*.ts` IS in extensions — transport stream is used for Live TV recordings tracked
|
||||
by Sonarr. Orphaned `.ts` recordings should be cleaned like any other orphaned episode.
|
||||
|
||||
---
|
||||
|
||||
### Radarr Cleanup Thresholds
|
||||
|
||||
```bash
|
||||
RADARR_ORPHAN_AGE=7
|
||||
RADARR_MAX_DELETE_GB=50
|
||||
RADARR_IMPORT_SCAN_TIMEOUT=600
|
||||
RADARR_VERSION_MAJOR=6
|
||||
RADARR_EXTENSIONS=("mkv" "mp4" "avi" "m4v" "wmv" "mov")
|
||||
RADARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo" "*.srt" "*.sub" "*.ass" "*.ssa")
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Arr Sync
|
||||
|
||||
```bash
|
||||
ARR_SYNC_ENABLED=true
|
||||
ARR_SYNC_BLOCKLIST=/boot/config/arr_sync_blocklist.tsv # tombstone file
|
||||
ARR_SYNC_CONNECT_TIMEOUT=10 # SSH connect timeout in seconds
|
||||
ARR_SYNC_API_TIMEOUT=60 # curl API call timeout in seconds
|
||||
DOCKER_APPDATA_BASE=/mnt/user/appdata
|
||||
ARR_SYNC_LIDARR_PORT=8686
|
||||
ARR_SYNC_SONARR_PORT=8989
|
||||
ARR_SYNC_RADARR_PORT=7878
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Arr Recovery
|
||||
|
||||
```bash
|
||||
ARR_IMPORT_RECOVERY_AGE=6 # hours — items newer than this are skipped
|
||||
SONARR_VERSION_MAJOR=4
|
||||
RADARR_VERSION_MAJOR=6
|
||||
LIDARR_VERSION_MAJOR=3
|
||||
ARR_RECOVERY_STATS=/boot/config/arr_recovery_stats.db # read by coffee report
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### TMDb / TVDB Removed
|
||||
|
||||
```bash
|
||||
RADARR_DROPPED_ADD_EXCLUSION=true # add removed movies to Radarr import exclusion
|
||||
SONARR_DROPPED_ADD_EXCLUSION=true # add removed series to Sonarr import exclusion
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Lidarr Missing Art
|
||||
|
||||
```bash
|
||||
FANART_API_KEY="your-fanart-tv-api-key"
|
||||
LASTFM_API_KEY="your-lastfm-api-key"
|
||||
LIDARR_ART_MIN_SIZE=5000 # minimum valid download size in bytes
|
||||
LIDARR_ART_MAX_PARALLEL=4 # concurrent background download jobs
|
||||
LIDARR_ART_RETRIES=2 # download retry attempts per image
|
||||
LIDARR_ART_SLEEP_BETWEEN=1 # seconds between fanart.tv API calls (rate limit)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Lidarr Discovery
|
||||
|
||||
```bash
|
||||
LIDARR_DISCOVERY_THRESHOLD=70 # minimum score for Stage 1 seeds and Stage 2 adds
|
||||
LIDARR_DISCOVERY_LOOKBACK_DAYS=7 # Emby play history window in days
|
||||
LIDARR_DISCOVERY_MIN_PLAYS=3 # min plays before an artist is evaluated as a seed
|
||||
LIDARR_DISCOVERY_MAX_ADDS=5 # max seeds (Stage 1) and max adds (Stage 2) per run
|
||||
LIDARR_DISCOVERY_USER_CAP_PCT=35 # max % any one user contributes to play weight
|
||||
LIDARR_DISCOVERY_REJECT_COOLDOWN=30 # days before re-evaluating a Stage 2 reject
|
||||
LIDARR_DISCOVERY_HISTORY="$DATA_DIR/lidarr_discovery_history.db"
|
||||
```
|
||||
|
||||
Requires `HOST*_LASTFM_API_KEY` in `host*.conf`.
|
||||
|
||||
---
|
||||
|
||||
### Radarr Discovery
|
||||
|
||||
```bash
|
||||
RADARR_DISCOVERY_THRESHOLD=52 # minimum score to add a candidate
|
||||
RADARR_DISCOVERY_LOOKBACK_DAYS=30 # Emby watch history window in days
|
||||
RADARR_DISCOVERY_MAX_SEEDS=5 # max seed movies from Stage 1
|
||||
RADARR_DISCOVERY_MAX_ADDS=5 # max movies to add per run
|
||||
RADARR_DISCOVERY_MIN_VOTE_COUNT=100 # min TMDB votes for a candidate
|
||||
RADARR_DISCOVERY_MIN_RATING=60 # min TMDB vote_average × 10 (60 = 6.0/10)
|
||||
RADARR_DISCOVERY_REJECT_COOLDOWN=60 # days before re-evaluating a rejected movie
|
||||
RADARR_DISCOVERY_SEED_LIBRARIES=("Movies") # Emby libraries to draw seeds from
|
||||
RADARR_DISCOVERY_HISTORY="$DATA_DIR/radarr_discovery_history.db"
|
||||
```
|
||||
|
||||
Requires `HOST*_TMDB_API_KEY` in `host*.conf`.
|
||||
|
||||
---
|
||||
|
||||
### Sonarr Discovery
|
||||
|
||||
```bash
|
||||
SONARR_DISCOVERY_THRESHOLD=52 # minimum score to add a candidate
|
||||
SONARR_DISCOVERY_LOOKBACK_DAYS=14 # Emby episode history window in days
|
||||
SONARR_DISCOVERY_MAX_SEEDS=5 # max seed series from Stage 1
|
||||
SONARR_DISCOVERY_MAX_ADDS=3 # max shows to add per run (TV is a larger commitment)
|
||||
SONARR_DISCOVERY_MIN_VOTE_COUNT=50 # min TMDB votes for a candidate
|
||||
SONARR_DISCOVERY_MIN_RATING=65 # min TMDB vote_average × 10 (65 = 6.5/10)
|
||||
SONARR_DISCOVERY_REJECT_COOLDOWN=60 # days before re-evaluating a rejected show
|
||||
SONARR_DISCOVERY_USER_EPISODE_CAP=8 # max episodes per user in seed scoring
|
||||
SONARR_DISCOVERY_MONITOR_MODE="all" # "all" = all seasons monitored; "future" = upcoming only
|
||||
SONARR_DISCOVERY_HISTORY="$DATA_DIR/sonarr_discovery_history.db"
|
||||
# SONARR_EMBY_LIBRARIES is shared with emby_to_sonarr_sync — see Orchestrator Job Order
|
||||
```
|
||||
|
||||
Requires `HOST*_TMDB_API_KEY` in `host*.conf`.
|
||||
|
||||
> **MONITOR_MODE note:** Use `"all"` (default) to have Sonarr search all existing seasons
|
||||
> after adding a show. `"future"` only marks upcoming seasons as monitored — shows where
|
||||
> all seasons have already aired will appear unmonitored and Sonarr will not search for them.
|
||||
|
||||
---
|
||||
|
||||
### Orchestrator Job Order
|
||||
|
||||
```bash
|
||||
MEDIA_MAINTENANCE_JOBS=(
|
||||
"Media/media_shares_permissions.sh" # 1. permissions — always first
|
||||
"Media/media_cleaner.sh anime" # 2. junk removal — before orphan scan
|
||||
"Media/media_cleaner.sh media" # 3.
|
||||
"Media/lidarr_cleanup.sh" # 4. arr cleanup — after permissions + clean
|
||||
"Media/sonarr_cleanup.sh" # 5.
|
||||
"Media/radarr_cleanup.sh" # 6.
|
||||
)
|
||||
# host*.conf
|
||||
HOST1_EMBY_URL="http://192.168.50.2:8096"
|
||||
HOST1_EMBY_API_KEY="..."
|
||||
HOST1_JELLYFIN_URL="" # empty = skip Jellyfin on this host
|
||||
HOST1_JELLYFIN_API_KEY=""
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ CONFIGURATION — host*.conf ━━━
|
||||
|
||||
### host1.conf
|
||||
|
||||
```bash
|
||||
# Shares this server applies permissions to
|
||||
HOST1_MEDIA_PERMISSION_SHARES=(
|
||||
@@ -321,224 +147,6 @@ HOST1_MEDIA_CLEAN_FOLDERS=(
|
||||
"/mnt/user/stand-up_comedy"
|
||||
"/mnt/user/Tv_Shows"
|
||||
)
|
||||
|
||||
# Arr connection details — must match arr settings exactly
|
||||
HOST1_LIDARR_URL="http://192.168.50.2:8686"
|
||||
HOST1_LIDARR_API_KEY="..."
|
||||
HOST1_LIDARR_MUSIC_ROOT="/mnt/user/Music"
|
||||
HOST1_LIDARR_PATH_MAP="" # container→host path translation if needed
|
||||
|
||||
HOST1_SONARR_URL="http://192.168.50.2:8989"
|
||||
HOST1_SONARR_API_KEY="..."
|
||||
HOST1_SONARR_TV_ROOT="/mnt/user/Tv_Shows"
|
||||
HOST1_SONARR_PATH_MAP=""
|
||||
|
||||
HOST1_RADARR_URL="http://192.168.50.2:7878"
|
||||
HOST1_RADARR_API_KEY="..."
|
||||
HOST1_RADARR_MOVIES_ROOT="/mnt/user/Movies"
|
||||
HOST1_RADARR_PATH_MAP=""
|
||||
|
||||
HOST1_EMBY_URL="http://192.168.50.2:8096"
|
||||
HOST1_EMBY_API_KEY="..."
|
||||
```
|
||||
|
||||
> **LIDARR/SONARR/RADARR_MUSIC/TV/MOVIES_ROOT must exactly match the Root Folder path in
|
||||
> the arr's own settings.** Arr UI → Settings → Media Management → Root Folders.
|
||||
> A mismatch means every file on disk looks untracked — all appear as orphans.
|
||||
> MAX_DELETE_GB is the only thing standing between a path mismatch and losing your library.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SAFE TESTING PROCEDURE ━━━
|
||||
|
||||
> **The arr cleanup scripts permanently delete files.** There is no recycle bin, no undo.
|
||||
> Follow this procedure on first use, after any root path change, after any API key change,
|
||||
> and after any significant arr library change.
|
||||
|
||||
### Step 1 — Dry Run With Full Logging
|
||||
|
||||
```bash
|
||||
lidarr_cleanup.sh --dry-run --log
|
||||
sonarr_cleanup.sh --dry-run --log
|
||||
radarr_cleanup.sh --dry-run --log
|
||||
```
|
||||
|
||||
### Step 2 — Review the Output
|
||||
|
||||
```
|
||||
Are TRACKED files the ones you expect?
|
||||
→ Known arr-managed files should show as TRACKED
|
||||
→ If they show as ORPHAN, the root path is wrong — STOP
|
||||
|
||||
Is the ORPHAN count reasonable?
|
||||
→ Healthy cleanup removes dozens to hundreds, not tens of thousands
|
||||
→ Large count = stop, investigate root path before proceeding
|
||||
|
||||
Are PROTECTED patterns working?
|
||||
→ Artwork (*.jpg) and subtitles (*.srt) must show as PROTECTED
|
||||
→ If they show as ORPHAN, check PROTECTED_PATTERNS config
|
||||
|
||||
Are RECENT files being correctly skipped?
|
||||
→ Files downloaded in the last 7 days should show as RECENT, not ORPHAN
|
||||
```
|
||||
|
||||
### Step 3 — Check Numbers if Something Looks Wrong
|
||||
|
||||
```bash
|
||||
# Root path mismatch? Compare these:
|
||||
# Lidarr UI: Settings → Media Management → Root Folders
|
||||
# Sonarr UI: Settings → Media Management → Root Folders
|
||||
# Radarr UI: Settings → Media Management → Root Folders
|
||||
# Must exactly match LIDARR_MUSIC_ROOT / SONARR_TV_ROOT / RADARR_MOVIES_ROOT
|
||||
|
||||
# Is the arr running?
|
||||
docker ps | grep -E "Lidarr|Sonarr|Radarr"
|
||||
|
||||
# Library scan not complete?
|
||||
# Trigger manual scan in arr UI and wait for completion
|
||||
```
|
||||
|
||||
### Step 4 — Run Live
|
||||
|
||||
```bash
|
||||
# Only after dry run review passes.
|
||||
lidarr_cleanup.sh
|
||||
sonarr_cleanup.sh
|
||||
radarr_cleanup.sh
|
||||
```
|
||||
|
||||
### Step 5 — Verify in Arr UI
|
||||
|
||||
```
|
||||
Library count — should not have dropped significantly
|
||||
healthy cleanup removes a small number, not a large percentage
|
||||
Missing files — check if any monitored content shows as missing
|
||||
Emby library — should show no ghost entries (notify_emby_scan handles this automatically)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ PROCEDURES ━━━
|
||||
|
||||
### Adding a New Arr
|
||||
|
||||
```bash
|
||||
# 1. Copy radarr_cleanup.sh as template
|
||||
cp radarr_cleanup.sh readarr_cleanup.sh
|
||||
|
||||
# 2. Replace RADARR_ prefix with READARR_ throughout
|
||||
# Update API endpoint, tracked file API path, extension list, protected patterns
|
||||
|
||||
# 3. Add to host*.conf
|
||||
HOST1_READARR_URL="http://192.168.50.2:8787"
|
||||
HOST1_READARR_API_KEY="your-api-key"
|
||||
HOST1_READARR_BOOKS_ROOT="/mnt/user/Books"
|
||||
|
||||
# 4. Add thresholds to master.conf
|
||||
READARR_ORPHAN_AGE=7
|
||||
READARR_MAX_DELETE_GB=50
|
||||
READARR_EXTENSIONS=("epub" "pdf" "mobi" "azw3" "cbz" "cbr")
|
||||
READARR_PROTECTED_PATTERNS=("*.jpg" "*.jpeg" "*.png" "*.nfo")
|
||||
|
||||
# 5. Add to MEDIA_MAINTENANCE_JOBS in master.conf
|
||||
MEDIA_MAINTENANCE_JOBS=(
|
||||
...existing jobs...
|
||||
"Media/readarr_cleanup.sh"
|
||||
)
|
||||
```
|
||||
|
||||
media_management.sh picks it up automatically. No orchestrator changes needed.
|
||||
Run `--dry-run --log` before scheduling.
|
||||
|
||||
---
|
||||
|
||||
### Managing the Arr Sync Blocklist
|
||||
|
||||
```bash
|
||||
# Add item to blocklist (removes from all arrs + tombstones the ID)
|
||||
arr_sync.sh --blocklist-add lidarr <musicbrainz-artist-id> "reason"
|
||||
arr_sync.sh --blocklist-add sonarr <tvdb-series-id> "reason"
|
||||
arr_sync.sh --blocklist-add radarr <tmdb-movie-id> "reason"
|
||||
|
||||
# Remove from blocklist (un-tombstones the ID — does NOT re-add to arrs)
|
||||
arr_sync.sh --blocklist-remove lidarr <id>
|
||||
|
||||
# View all blocklisted IDs
|
||||
arr_sync.sh --blocklist-list
|
||||
```
|
||||
|
||||
`--blocklist-add` is the only destructive operation — it simultaneously:
|
||||
1. Writes the tombstone entry to the blocklist TSV file
|
||||
2. Deletes the item from the local arr API (no file deletion)
|
||||
3. SSHes each remote node and deletes from their arr API
|
||||
|
||||
Files become orphans on all nodes — arr_cleanup removes them on the next run.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ TROUBLESHOOTING ━━━
|
||||
|
||||
### Arr Cleanup Deleting Files It Shouldn't
|
||||
|
||||
```
|
||||
1. Check the protected patterns — artwork and subtitles must be listed
|
||||
LIDARR_PROTECTED_PATTERNS / SONARR_PROTECTED_PATTERNS / RADARR_PROTECTED_PATTERNS
|
||||
|
||||
2. Check the root path matches arr settings exactly
|
||||
Run: lidarr_cleanup.sh --status (shows configured root path)
|
||||
Compare: Lidarr UI → Settings → Media Management → Root Folders
|
||||
|
||||
3. Check if files are truly orphaned
|
||||
Run: lidarr_cleanup.sh --dry-run --log
|
||||
Look for the specific file — verify it shows ORPHAN, not PROTECTED or TRACKED
|
||||
```
|
||||
|
||||
### Arr Cleanup Aborting at Safety Layer 6 (Tracked Count Drop)
|
||||
|
||||
```
|
||||
API returned far fewer tracked files than last run.
|
||||
Possible causes:
|
||||
- Arr database was recently rebuilt from scratch
|
||||
- Large manual library removal
|
||||
- Path map mismatch after arr migration
|
||||
|
||||
If intentional (library intentionally reduced):
|
||||
Delete LIDARR_TRACKED_COUNT_FILE to reset the baseline
|
||||
Run cleanup once — it will establish a new baseline
|
||||
|
||||
If unintentional:
|
||||
Investigate before proceeding — the arr may have a problem
|
||||
```
|
||||
|
||||
### Arr Sync Not Picking Up New Content
|
||||
|
||||
```
|
||||
Is ARR_SYNC_ENABLED=true in master.conf?
|
||||
|
||||
Can this host SSH to the remote without password?
|
||||
→ ssh -i [SSH_KEY] root@[remote-tailscale-ip] "hostname"
|
||||
|
||||
Is the arr accessible on the remote?
|
||||
→ arr_sync.sh --status (shows each node's arr reachability)
|
||||
→ arr_sync.sh --log (verbose output per-node, per-arr)
|
||||
|
||||
Is the item in the blocklist?
|
||||
→ arr_sync.sh --blocklist-list
|
||||
```
|
||||
|
||||
### Emby Still Showing Ghost Entries After Cleanup
|
||||
|
||||
```
|
||||
notify_emby_scan() is called automatically after every arr cleanup deletion.
|
||||
If ghosts persist:
|
||||
1. Is Emby's API responding?
|
||||
curl -s "http://[emby-ip]:8096/System/Info/Public"
|
||||
|
||||
2. Is EMBY_URL / EMBY_API_KEY correct in host*.conf?
|
||||
Run: sonarr_cleanup.sh --status (shows Emby config)
|
||||
|
||||
3. Trigger manually in Emby:
|
||||
Library → Manage Library → Clean Missing Files
|
||||
```
|
||||
|
||||
---
|
||||
@@ -547,16 +155,11 @@ If ghosts persist:
|
||||
|
||||
All scripts have two output levels controlled by `--log`.
|
||||
|
||||
Without `--log`, each script processes silently and always concludes with a summary
|
||||
block: identity, duration, counts (files removed, items added, arrs cleaned), and a
|
||||
status line. Warnings and errors are always visible.
|
||||
Without `--log`, each script processes silently and concludes with a summary block.
|
||||
Warnings and errors are always visible.
|
||||
|
||||
With `--log`, per-item detail appears: individual titles being processed, API query
|
||||
progress, per-node sync results, and per-file examination output. Use this when
|
||||
debugging unexpected results or validating configuration before the first scheduled run.
|
||||
|
||||
Dry-run output follows the same tiers — `--dry-run` alone shows the summary of what
|
||||
would happen; `--dry-run --log` shows the full per-item preview list.
|
||||
With `--log`, per-item detail appears — individual shares being processed, file counts,
|
||||
per-server sync results.
|
||||
|
||||
---
|
||||
|
||||
@@ -564,18 +167,18 @@ would happen; `--dry-run --log` shows the full per-item preview list.
|
||||
|
||||
### media_shares_permissions.sh
|
||||
|
||||
`media_shares_permissions.sh`
|
||||
`media_shares_permissions.sh`
|
||||
Apply correct ownership and permissions to all configured media shares. Safe to run
|
||||
manually at any time — idempotent, only changes what's wrong.
|
||||
|
||||
`media_shares_permissions.sh --dry-run`
|
||||
`media_shares_permissions.sh --dry-run`
|
||||
Show how many files and directories would be corrected per share. If unexpectedly large,
|
||||
check container PUID/PGID settings first (PUID=99 PGID=100).
|
||||
|
||||
`media_shares_permissions.sh --status`
|
||||
`media_shares_permissions.sh --status`
|
||||
Show configured share list and ownership of the share roots.
|
||||
|
||||
`media_shares_permissions.sh --log`
|
||||
`media_shares_permissions.sh --log`
|
||||
Show ownership correction count per share and per file (verbose).
|
||||
|
||||
> On large libraries this runs 20-30 minutes. This is expected — millions of files with
|
||||
@@ -585,183 +188,51 @@ Show ownership correction count per share and per file (verbose).
|
||||
|
||||
### media_cleaner.sh
|
||||
|
||||
`media_cleaner.sh anime`
|
||||
`media_cleaner.sh anime`
|
||||
Remove junk files from anime share folders using ANIME_FILE_PATTERNS.
|
||||
|
||||
`media_cleaner.sh media`
|
||||
`media_cleaner.sh media`
|
||||
Remove junk files from media share folders using MEDIA_FILE_PATTERNS.
|
||||
|
||||
`media_cleaner.sh [profile] --dry-run`
|
||||
`media_cleaner.sh [profile] --dry-run`
|
||||
Show what would be deleted without removing anything. Always run first when adding new
|
||||
patterns or folders.
|
||||
|
||||
`media_cleaner.sh [profile] --status`
|
||||
`media_cleaner.sh [profile] --status`
|
||||
Show folder list and file patterns for the profile.
|
||||
|
||||
`media_cleaner.sh [profile] --log`
|
||||
`media_cleaner.sh [profile] --log`
|
||||
Show every file examined, not just those removed.
|
||||
|
||||
---
|
||||
|
||||
### lidarr_cleanup.sh / sonarr_cleanup.sh / radarr_cleanup.sh
|
||||
### play_state_sync.sh
|
||||
|
||||
`[script] --dry-run --log`
|
||||
Preview every classification decision. **Always run this first.** See Safe Testing Procedure.
|
||||
`play_state_sync.sh`
|
||||
Sync played/unplayed state, resume positions and favourites across **every** configured Emby
|
||||
and Jellyfin server — not just local→remote. Servers are discovered from every
|
||||
`HOST*_TRANSCODE_SERVERS` entry, with remote hosts' localhost URLs rewritten to their
|
||||
Tailscale IP. Newest `LastPlayedDate` wins; state only ever moves forward, never clears.
|
||||
|
||||
`[script]`
|
||||
Live run — deletes confirmed orphans and junk, triggers Emby clean.
|
||||
No age filter — every matched item is considered on every run. The **change probe** is what
|
||||
keeps that cheap: the raw API responses are hashed and compared against the fingerprint from
|
||||
the last successful run, and per-item processing is skipped entirely when nothing moved.
|
||||
Fetches still happen every run, so nothing can be missed by the probe.
|
||||
|
||||
`[script] --log`
|
||||
Live run with verbose per-file output.
|
||||
`play_state_sync.sh --full`
|
||||
Bypass the change probe and force the full per-item comparison even when the fingerprint
|
||||
matches. Use after a new Emby install or database restore, or when debugging a sync that
|
||||
appears to be skipping work it should be doing.
|
||||
|
||||
`[script] --status`
|
||||
Show configuration, API status, tracked file count, and last run stats.
|
||||
`play_state_sync.sh --wait`
|
||||
Wait for an in-progress run instead of exiting. For manual runs that would otherwise be
|
||||
skipped by the scheduled every-30-minute pass. Used by `fallback.sh` during handback.
|
||||
|
||||
`[script] --i-know-what-im-doing`
|
||||
Bypass the MAX_DELETE_GB size threshold. Required when deletion exceeds the configured
|
||||
limit. Long flag name is intentional — cannot be added accidentally.
|
||||
`play_state_sync.sh --dry-run`
|
||||
Show what would be synced without writing any state.
|
||||
|
||||
`[script] --skip-strike-list`
|
||||
Bypass the ORPHAN_AGE age check. Deletes RECENT files too — files that are under the
|
||||
age threshold. Use when you know recent downloads are actually orphans.
|
||||
`play_state_sync.sh --status`
|
||||
Show configured servers, reachability, and user counts.
|
||||
|
||||
`[script] --i-know-what-im-doing --skip-strike-list`
|
||||
**NUCLEAR MODE** — age check and size threshold both bypassed. Deletes on first pass.
|
||||
Use when you want a clean one-pass wipe of everything the arr doesn't track.
|
||||
No recovery possible after deletion.
|
||||
|
||||
---
|
||||
|
||||
### arrs_failed_stalled_recovery.sh
|
||||
|
||||
`arrs_failed_stalled_recovery.sh`
|
||||
Check all configured arrs for failed imports and stalled downloads. Blocklist + remove +
|
||||
re-search for each problem item.
|
||||
|
||||
`arrs_failed_stalled_recovery.sh --dry-run`
|
||||
Show what would be actioned per arr without making any changes.
|
||||
|
||||
`arrs_failed_stalled_recovery.sh --status`
|
||||
Show configuration, arr reachability, and last recovery stats.
|
||||
|
||||
`arrs_failed_stalled_recovery.sh --log`
|
||||
Verbose output per item per arr.
|
||||
|
||||
---
|
||||
|
||||
### arr_sync.sh
|
||||
|
||||
`arr_sync.sh`
|
||||
Sync all arr types across all configured nodes.
|
||||
|
||||
`arr_sync.sh --dry-run`
|
||||
Show what would be added/removed on each node without making changes.
|
||||
|
||||
`arr_sync.sh --status`
|
||||
Show node configuration, arr reachability, and blocklist count.
|
||||
|
||||
`arr_sync.sh --log`
|
||||
Verbose per-node, per-arr output.
|
||||
|
||||
`arr_sync.sh --blocklist-add [arr] [id] "[reason]"`
|
||||
Remove item from all arrs and tombstone the ID. See Procedures above.
|
||||
|
||||
`arr_sync.sh --blocklist-remove [arr] [id]`
|
||||
Remove tombstone — does NOT re-add item to arrs.
|
||||
|
||||
`arr_sync.sh --blocklist-list`
|
||||
Show all tombstoned IDs.
|
||||
|
||||
---
|
||||
|
||||
### radarr_tmdb_removed.sh / sonarr_tvdb_removed.sh
|
||||
|
||||
`[script]`
|
||||
Remove records for entries with status="deleted" (dropped from upstream database).
|
||||
Files are kept. Import exclusion is added.
|
||||
|
||||
`[script] --delete-files`
|
||||
Also delete associated files from disk. Most dropped entries have no files — they were
|
||||
announced movies/series that were never downloaded.
|
||||
|
||||
`[script] --dry-run`
|
||||
Preview what would be removed without making changes.
|
||||
|
||||
`[script] --status`
|
||||
Show arr connection status and current count of dropped entries.
|
||||
|
||||
`[script] --log`
|
||||
Verbose per-entry output.
|
||||
|
||||
---
|
||||
|
||||
### lidarr_missing_art.sh
|
||||
|
||||
`lidarr_missing_art.sh`
|
||||
Fetch all missing album and artist artwork from fanart.tv and fallback sources.
|
||||
Never overwrites existing files.
|
||||
|
||||
`lidarr_missing_art.sh --dry-run`
|
||||
Show what would be downloaded without writing any files.
|
||||
|
||||
`lidarr_missing_art.sh --status`
|
||||
Show configuration and API key status.
|
||||
|
||||
`lidarr_missing_art.sh --log`
|
||||
Verbose per-album, per-artist output.
|
||||
|
||||
---
|
||||
|
||||
### playback_aware_lidarr_discovery.sh
|
||||
|
||||
`playback_aware_lidarr_discovery.sh`
|
||||
Score Emby play history, run Last.fm getSimilar on top artists, add candidates above
|
||||
threshold to Lidarr. Triggers ArtistSearch immediately after each successful add.
|
||||
|
||||
`playback_aware_lidarr_discovery.sh --dry-run`
|
||||
Score and rank all Stage 1 seeds and Stage 2 candidates. No Lidarr API calls. No writes
|
||||
to history file. Shows exactly what would be added and at what score.
|
||||
|
||||
`playback_aware_lidarr_discovery.sh --status`
|
||||
Show config values, history file path and size, and API key status.
|
||||
|
||||
`playback_aware_lidarr_discovery.sh --log`
|
||||
Verbose per-artist scoring output for both stages.
|
||||
|
||||
---
|
||||
|
||||
### playback_aware_radarr_discovery.sh
|
||||
|
||||
`playback_aware_radarr_discovery.sh`
|
||||
Score recently watched Emby movies, run TMDB recommendations on seeds, add candidates
|
||||
above threshold to Radarr. Triggers MoviesSearch immediately after each successful add.
|
||||
|
||||
`playback_aware_radarr_discovery.sh --dry-run`
|
||||
Score and rank all Stage 1 seeds and Stage 2 candidates. No Radarr API calls. No writes
|
||||
to history file.
|
||||
|
||||
`playback_aware_radarr_discovery.sh --status`
|
||||
Show config values, history file path and size, and API key status.
|
||||
|
||||
`playback_aware_radarr_discovery.sh --log`
|
||||
Verbose per-movie scoring output for both stages.
|
||||
|
||||
---
|
||||
|
||||
### playback_aware_sonarr_discovery.sh
|
||||
|
||||
`playback_aware_sonarr_discovery.sh`
|
||||
Score recently watched Emby series (weighted by user diversity), run TMDB TV
|
||||
recommendations on seeds, add candidates above threshold to Sonarr. Triggers SeriesSearch
|
||||
immediately after each successful add.
|
||||
|
||||
`playback_aware_sonarr_discovery.sh --dry-run`
|
||||
Score and rank all Stage 1 seeds and Stage 2 candidates. No Sonarr API calls. No writes
|
||||
to history file.
|
||||
|
||||
`playback_aware_sonarr_discovery.sh --status`
|
||||
Show config values, history file path and size, and API key status.
|
||||
|
||||
`playback_aware_sonarr_discovery.sh --log`
|
||||
Verbose per-series scoring output — shows user diversity, recency, and volume scores per
|
||||
seed; breadth, rating, and votes scores per candidate.
|
||||
`play_state_sync.sh --log`
|
||||
Verbose output — show each item comparison.
|
||||
|
||||
+68
-161
@@ -1,149 +1,73 @@
|
||||
# ━━━━━ MEDIA ━━━━━
|
||||
|
||||
Library health, consistency, sync, and behavior-driven discovery for a multi-server arr
|
||||
stack. Correct permissions so arrs can manage files. Junk removal so orphan detection
|
||||
isn't confused by scene debris. Library sync so every node tracks the same content.
|
||||
Orphan cleanup against live arr APIs so deleted content actually leaves disk. Emby
|
||||
notified automatically after every deletion. Weekly discovery adds new music, movies,
|
||||
and TV shows based on what you actually play — no manual browsing required.
|
||||
Foundation-level media library management — permissions, junk removal, and play state
|
||||
sync. These three scripts run before and independently of arr stack operations.
|
||||
|
||||
> **These scripts permanently delete files.** The arr cleanup scripts are protected by
|
||||
> multiple safety layers that must all pass before anything is touched — but dry runs and
|
||||
> log review are still the right first step on any new system or after any configuration
|
||||
> change. The testing procedure in Manual-Media.md exists for a reason.
|
||||
For arr stack scripts (orphan cleanup, release fixer, sync, discovery, webhooks) see
|
||||
`Arrs_Stack/README-Arrs_Stack.md`.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE PROBLEM THAT BUILT THIS ━━━
|
||||
## ━━━ THE PROBLEMS THAT BUILT THIS ━━━
|
||||
|
||||
**Files Owned by Root That Arrs Can't Touch**
|
||||
**Files Owned by Root That Arrs Can't Touch**
|
||||
Download clients without explicit PUID/PGID write files owned by root. Arrs running as
|
||||
`nobody:users` cannot rename, move, or delete them. Import fails. Upgrade attempts fail.
|
||||
The failure is subtle — arr shows the file as managed but can't touch it. You only
|
||||
discover this when an upgrade is requested and the old version refuses to delete.
|
||||
discover this when an upgrade is requested and the old version refuses to delete.
|
||||
The fix: `media_shares_permissions.sh` applies correct ownership daily. Even if a
|
||||
container is misconfigured, the window is at most 24 hours.
|
||||
|
||||
**Scene Junk Confusing Orphan Detection**
|
||||
**Scene Junk Confusing Orphan Detection**
|
||||
Scene releases include `.sfv`, `.nfo`, `.rar`, `.sample` files alongside the actual media.
|
||||
After extraction and import these are worthless — but they're not tracked by any arr.
|
||||
They look like orphans. Processing them as orphans means the cleanup output is full of
|
||||
noise, making it hard to spot actual orphaned media.
|
||||
noise, making it hard to spot actual orphaned media.
|
||||
The fix: `media_cleaner.sh` runs before any arr cleanup and removes all known junk
|
||||
patterns first. By the time arr cleanup runs, every untracked file is actual media.
|
||||
|
||||
**Deleted Shows and Removed Albums Still on Disk**
|
||||
When you remove a series from Sonarr and the delete command fails — permission issue,
|
||||
container wasn't running, path mismatch — the files stay permanently. Over years on an
|
||||
active library this accumulates significantly.
|
||||
The fix: arr cleanup scripts query the live API for every tracked file path, walk the
|
||||
disk, and delete anything absent from the API response that's old enough to be past the
|
||||
import window.
|
||||
|
||||
**Emby Showing Ghost Entries After Cleanup**
|
||||
After arr cleanup deletes files, Emby still shows them until its next scheduled scan —
|
||||
potentially hours later. Users see broken entries that produce "file not found" errors.
|
||||
The fix: `notify_emby_scan()` is called automatically after every deletion. Triggers
|
||||
Emby's "Clean Missing Files" task immediately.
|
||||
|
||||
**No Safety Net on Deletion Size**
|
||||
A misconfigured root path — pointing cleanup at the wrong directory — means the API
|
||||
returns zero tracked files for a root that actually contains thousands. Every file walks
|
||||
as an orphan. Everything gets deleted. This is the catastrophic failure mode.
|
||||
The fix: `LIDARR/SONARR/RADARR_MAX_DELETE_GB` — if total deletion size exceeds the
|
||||
limit, the script stops and requires `--i-know-what-im-doing` to proceed. The flag name
|
||||
is long and annoying by design. It cannot be added by accident.
|
||||
**Watch State Diverging Across Servers**
|
||||
With two Emby servers, played status and resume positions diverge — a film marked watched
|
||||
on HOST1 shows as unwatched on HOST2. Two users on different servers get different
|
||||
continue-watching rows.
|
||||
The fix: `play_state_sync.sh` syncs watched/played state and resume positions every 30
|
||||
minutes. Newest timestamp wins. Both servers always reflect the same play history.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ WHAT THIS FOLDER DOES ━━━
|
||||
|
||||
Scripts divide into five functional areas:
|
||||
|
||||
**Library Foundation**
|
||||
**Library Foundation**
|
||||
`media_shares_permissions.sh` — normalize ownership and permissions daily. Runs first in
|
||||
every maintenance window because arr cleanup depends on correct ownership to delete files.
|
||||
|
||||
**Junk Removal**
|
||||
**Junk Removal**
|
||||
`media_cleaner.sh` — remove scene debris and tool artifacts before orphan scan. Runs
|
||||
second, before any arr cleanup, so orphan detection only encounters actual media files.
|
||||
before any arr cleanup so orphan detection only encounters actual media files.
|
||||
|
||||
**Library Sync**
|
||||
`arr_sync.sh` — full-mesh arr library sync across all nodes. Every node syncs with every
|
||||
other — union model, no hierarchy. Once arrs agree on what to track, rsync spreads the
|
||||
actual files. The architectural shift from file-first sync to arr-first sync.
|
||||
|
||||
**Orphan Cleanup**
|
||||
`lidarr_cleanup.sh`, `sonarr_cleanup.sh`, `radarr_cleanup.sh` — API-verified orphan
|
||||
removal. Five classification categories (TRACKED/PROTECTED/ORPHAN/JUNK/RECENT), seven
|
||||
safety layers, automatic Emby notification after deletion.
|
||||
|
||||
**Database Hygiene**
|
||||
`radarr_tmdb_removed.sh`, `sonarr_tvdb_removed.sh` — remove entries that upstream
|
||||
databases have dropped (TMDb/TVDB status="deleted"). These generate health warnings in
|
||||
arrs and can never be monitored or downloaded. Most are announced-but-never-released
|
||||
entries. Files are kept by default — most have none.
|
||||
|
||||
**Library Enrichment**
|
||||
`arrs_failed_stalled_recovery.sh` — detect and recover failed imports and stalled
|
||||
downloads across all arrs. Blocklists the bad release and triggers a re-search — hands-
|
||||
free overnight recovery.
|
||||
`lidarr_missing_art.sh` — fetch missing album and artist artwork from fanart.tv and
|
||||
fallback sources. Never overwrites existing files.
|
||||
|
||||
**Discovery**
|
||||
`playback_aware_lidarr_discovery.sh` — behavior-driven music discovery. Scores your
|
||||
Emby play history, runs Last.fm getSimilar on top artists, adds the best matches to
|
||||
Lidarr. 0–5 meaningful adds per week.
|
||||
`playback_aware_radarr_discovery.sh` — behavior-driven movie discovery. Scores recently
|
||||
watched movies, runs TMDB recommendations on seeds, adds top candidates to Radarr.
|
||||
`playback_aware_sonarr_discovery.sh` — behavior-driven TV discovery. Scores recently
|
||||
watched series weighted by user diversity, runs TMDB TV recommendations on seeds, adds
|
||||
top shows to Sonarr. Multi-user design: one person binge-watching does not dominate seeds.
|
||||
**Play State Sync**
|
||||
`play_state_sync.sh` — syncs watched/played state and resume positions across all
|
||||
configured Emby and Jellyfin servers. Newest timestamp wins. Runs every 30 minutes
|
||||
via critical_sync_maintenance.sh.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ EXECUTION ORDER ━━━
|
||||
|
||||
Scripts run in multiple contexts — not all are part of the daily maintenance window:
|
||||
|
||||
**Daily via `media_management.sh` (MEDIA_MAINTENANCE_JOBS):**
|
||||
**Daily via `daily_sync_maintenance.sh` (DAILY_MAINTENANCE_SCRIPTS — runs first):**
|
||||
|
||||
```
|
||||
1. media_shares_permissions.sh — permissions first — arr cleanup depends on this
|
||||
2. media_cleaner.sh anime — junk before orphan scan
|
||||
3. media_cleaner.sh media
|
||||
4. lidarr_cleanup.sh — after permissions + clean
|
||||
5. sonarr_cleanup.sh
|
||||
6. radarr_cleanup.sh
|
||||
```
|
||||
|
||||
**Weekly arr sync (before rsync in weekly_sync_maintenance.sh):**
|
||||
These complete before any Arrs_Stack/ scripts run.
|
||||
|
||||
**Every 30 min via `critical_sync_maintenance.sh` (CRITICAL_MAINTENANCE_SCRIPTS):**
|
||||
|
||||
```
|
||||
arr_sync.sh — arrs agree on library → rsync then spreads the files
|
||||
```
|
||||
|
||||
**Daily recovery (separate schedule — 5am or every 6hr):**
|
||||
|
||||
```
|
||||
arrs_failed_stalled_recovery.sh
|
||||
```
|
||||
|
||||
**Weekly discovery (WEEKLY_MAINTENANCE_SCRIPTS in master.conf):**
|
||||
|
||||
```
|
||||
playback_aware_lidarr_discovery.sh — score play history → Last.fm similar → add to Lidarr
|
||||
playback_aware_radarr_discovery.sh — score watch history → TMDB recommendations → add to Radarr
|
||||
playback_aware_sonarr_discovery.sh — score episode history → TMDB TV recommendations → add to Sonarr
|
||||
```
|
||||
|
||||
**Ad-hoc or separate schedule:**
|
||||
|
||||
```
|
||||
lidarr_missing_art.sh — fetch missing artwork
|
||||
radarr_tmdb_removed.sh — weekly cleanup of TMDb-dropped entries
|
||||
sonarr_tvdb_removed.sh — weekly cleanup of TVDB-dropped entries
|
||||
play_state_sync.sh — sync watched/resume state across Emby + Jellyfin
|
||||
```
|
||||
|
||||
**Why permissions before everything else:** arr cleanup needs `nobody:users` ownership to
|
||||
@@ -154,82 +78,65 @@ but stays on disk.
|
||||
orphans. Removing them first means orphan detection only finds actual media. Cleaner
|
||||
output, more accurate detection.
|
||||
|
||||
**Why arr_sync before rsync:** once arrs agree on what to track, rsync spreads the actual
|
||||
files. An upgrade on one node — new tracked path, old path no longer in API — gets
|
||||
cleaned by arr_cleanup on all nodes after the next sync cycle.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ HOST AWARENESS ━━━
|
||||
|
||||
Scripts run on both servers via `detect_hosts()`, which aliases all `HOST*_` prefixed vars
|
||||
to their unprefixed names at runtime. No manual `HOST1`/`HOST2` comparisons exist in any script.
|
||||
|
||||
arr_sync.sh keeps all arr databases in bidirectional union — either server can download
|
||||
to any share. Arr cleanup uses the union model: a file is only an orphan if neither arr
|
||||
on either server has it indexed. Arr scripts check the aliased URL — if empty (arr not
|
||||
configured on this host), they exit cleanly with no action.
|
||||
to their unprefixed names at runtime.
|
||||
|
||||
Permissions and cleaner scripts run locally against each server's own shares, defined
|
||||
in `HOST*_MEDIA_PERMISSION_SHARES` and `HOST*_MEDIA_CLEAN_FOLDERS` in host*.conf.
|
||||
|
||||
`play_state_sync.sh` reads both servers' Emby/Jellyfin endpoints from host*.conf and
|
||||
syncs between them.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ SCRIPTS IN THIS FOLDER ━━━
|
||||
|
||||
| Script | Role | When It Runs |
|
||||
|--------|------|--------------|
|
||||
| `media_shares_permissions.sh` | Apply `nobody:users` ownership + correct permissions to all media shares | Daily via `media_management.sh` |
|
||||
| `media_cleaner.sh` | Remove junk files (two profiles: `anime` + `media`) | Daily via `media_management.sh` |
|
||||
| `arr_sync.sh` | Full-mesh arr library sync — all nodes track the same content | Weekly before rsync |
|
||||
| `lidarr_cleanup.sh` | Delete orphaned music files not tracked by Lidarr | Daily via `media_management.sh` |
|
||||
| `sonarr_cleanup.sh` | Delete orphaned TV files not tracked by Sonarr | Daily via `media_management.sh` |
|
||||
| `radarr_cleanup.sh` | Delete orphaned movie files not tracked by Radarr | Daily via `media_management.sh` |
|
||||
| `arrs_failed_stalled_recovery.sh` | Auto-recover failed imports and stalled downloads | Daily (5am or every 6hr) |
|
||||
| `lidarr_missing_art.sh` | Fetch missing album and artist artwork | Ad-hoc or separate schedule |
|
||||
| `radarr_tmdb_removed.sh` | Remove movies dropped from TMDb | Ad-hoc or weekly |
|
||||
| `sonarr_tvdb_removed.sh` | Remove series dropped from TVDB | Ad-hoc or weekly |
|
||||
| `playback_aware_lidarr_discovery.sh` | Behavior-driven music discovery — Emby plays → Last.fm similar → Lidarr | Weekly via `weekly_sync_maintenance.sh` |
|
||||
| `playback_aware_radarr_discovery.sh` | Behavior-driven movie discovery — Emby watches → TMDB recommendations → Radarr | Weekly via `weekly_sync_maintenance.sh` |
|
||||
| `playback_aware_sonarr_discovery.sh` | Behavior-driven TV discovery — Emby episodes → TMDB TV recommendations → Sonarr | Weekly via `weekly_sync_maintenance.sh` |
|
||||
| `media_shares_permissions.sh` | Apply `nobody:users` ownership + correct permissions to all media shares | Daily — runs first |
|
||||
| `media_cleaner.sh` | Remove junk files (two profiles: `anime` + `media`) | Daily — runs before arr cleanup |
|
||||
| `play_state_sync.sh` | Sync watched/played state + resume positions across Emby + Jellyfin | Every 30 min |
|
||||
|
||||
---
|
||||
|
||||
## ━━━ HOW THE SCRIPTS RELATE ━━━
|
||||
## ━━━ THE ctime INVARIANT — READ BEFORE CHANGING PERMISSIONS ━━━
|
||||
|
||||
`media_shares_permissions.sh` applies every pass **conditionally** — it touches only entries
|
||||
whose owner or mode is actually wrong. That is not an optimisation, and it must stay that way.
|
||||
|
||||
`chown` and `chmod` rewrite an inode's ctime **even when the value does not change**. A
|
||||
blanket pass would therefore restamp every file in the library every night.
|
||||
|
||||
The arr cleanup scripts (`sonarr_cleanup.sh`, `radarr_cleanup.sh`, `lidarr_cleanup.sh`) gate
|
||||
orphan deletion on ctime. mtime cannot substitute: an import preserves the release's original
|
||||
timestamp, so mtime says nothing about when a file arrived here. Measured 2026-07-27 — of 400
|
||||
files imported that week, **all 400 had mtimes over 7 days old, one of them 9613 days.**
|
||||
|
||||
So:
|
||||
|
||||
```
|
||||
Weekly arr sync (before rsync):
|
||||
arr_sync.sh ──────────────── syncs tracked IDs across all nodes
|
||||
│ union model: any node adds → all nodes get it
|
||||
│
|
||||
└── then rsync spreads the actual files to all nodes
|
||||
└── then arr_cleanup removes orphans on all nodes (old paths, removed content)
|
||||
|
||||
Daily maintenance window (media_management.sh):
|
||||
media_shares_permissions.sh
|
||||
│ (permissions correct — arr can now delete files)
|
||||
▼
|
||||
media_cleaner.sh (anime + media)
|
||||
│ (junk removed — orphan scan finds only actual media)
|
||||
▼
|
||||
lidarr_cleanup.sh ──────────── queries Lidarr API → walks /Music → deletes orphans
|
||||
sonarr_cleanup.sh ──────────── queries Sonarr API → walks /Tv_Shows → deletes orphans
|
||||
radarr_cleanup.sh ──────────── queries Radarr API → walks /Movies → deletes orphans
|
||||
│
|
||||
└── each cleanup → notify_emby_scan() → Emby removes ghost entries
|
||||
|
||||
Daily recovery:
|
||||
arrs_failed_stalled_recovery.sh ── importFailed/stalled → blocklist → re-search
|
||||
|
||||
Weekly discovery (WEEKLY_MAINTENANCE_SCRIPTS):
|
||||
playback_aware_lidarr_discovery.sh ─ Emby plays → Last.fm similar → top candidates → Lidarr
|
||||
playback_aware_radarr_discovery.sh ─ Emby watches → TMDB recommendations → top candidates → Radarr
|
||||
playback_aware_sonarr_discovery.sh ─ Emby episodes → TMDB TV recommendations → top candidates → Sonarr
|
||||
│
|
||||
└── each discovery script fires arr search immediately after successful add
|
||||
|
||||
Ad-hoc enrichment:
|
||||
lidarr_missing_art.sh ─────── discovers missing artwork → fetches from fanart.tv
|
||||
radarr_tmdb_removed.sh ────── status="deleted" → remove from Radarr + add exclusion
|
||||
sonarr_tvdb_removed.sh ────── status="deleted" → remove from Sonarr + add exclusion
|
||||
blanket chown/chmod → every ctime resets to today
|
||||
→ no file ever appears older than *_ORPHAN_AGE
|
||||
→ orphan collection silently stops
|
||||
→ nothing errors, nothing warns, disk just fills
|
||||
```
|
||||
|
||||
**The failure is invisible.** No script fails, no notification fires. The only symptom is
|
||||
orphans quietly accumulating until a pool fills — which is exactly how the 755 GB / 89%-full
|
||||
cache pool incident happened.
|
||||
|
||||
Two rules follow, and both are load-bearing:
|
||||
|
||||
1. **`media_shares_permissions.sh` passes stay conditional.** Making any of them unconditional
|
||||
breaks orphan collection ecosystem-wide.
|
||||
2. **`Tools/bulk_permissions_repair.sh` is unconditional by design** — it exists to repair
|
||||
known-wrong paths where correctness beats preserving a clock. That is precisely why it is a
|
||||
manual, targeted tool and not scheduled. Pointing it at a whole media root pauses orphan
|
||||
collection there for `*_ORPHAN_AGE` days.
|
||||
|
||||
Both scripts' headers carry this warning too. If you are reading this because you are about to
|
||||
"simplify" the permissions job, this is the thing that breaks.
|
||||
|
||||
@@ -1,461 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================= Arrs Failed / Stalled Recovery =====================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Detect and recover failed imports and stalled downloads across Sonarr, Radarr,
|
||||
# and Lidarr. Blocklists the bad release and triggers a re-search — hands-free
|
||||
# overnight recovery.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Four problem types detected from the arr queue API:
|
||||
# importFailed — downloaded but arr couldn't import the file
|
||||
# importPending — downloaded, stuck waiting to import (will not self-resolve)
|
||||
# error status — serious failure not covered by the above two states
|
||||
# stalled — download stuck with no connections or no progress
|
||||
#
|
||||
# Never touches items with state "downloading" or "imported" — safe to run anytime.
|
||||
# Items newer than ARR_IMPORT_RECOVERY_AGE are skipped — gives arr time to retry first.
|
||||
#
|
||||
# Per problem item (3-step response):
|
||||
# 1. Blocklist the release — prevents re-grabbing the same bad release
|
||||
# 2. Remove from queue — cleans up the failed item
|
||||
# 3. Trigger new search — finds a different release automatically
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# acquire_lock — prevents concurrent runs overlapping
|
||||
# jq validation — exits if jq not installed (required for JSON parsing)
|
||||
# API pre-flight — checks each arr is reachable before querying queue
|
||||
# Version check — check_arr_version() verifies running arr matches master.conf major
|
||||
# version; exits rather than silently misoperating after upgrade
|
||||
# Age threshold — skips items newer than ARR_IMPORT_RECOVERY_AGE (default 6hr)
|
||||
# Silent by default — only problems produce output, clean arrs stay silent
|
||||
#
|
||||
# API version mapping (endpoint paths differ from major version labels):
|
||||
# Sonarr v4 → /api/v3/ (v3 endpoint retained in v4)
|
||||
# Radarr v6 → /api/v3/ (v3 endpoint retained in v6)
|
||||
# Lidarr v3 → /api/v1/ (different from Sonarr/Radarr)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ARR_RECOVERY_STATS — stats file written after each run (read by coffee report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_SONARR_URL / HOST*_SONARR_API_KEY / HOST*_SONARR_RECOVERY
|
||||
# HOST*_RADARR_URL / HOST*_RADARR_API_KEY / HOST*_RADARR_RECOVERY
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY / HOST1_LIDARR_RECOVERY
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# ARR_IMPORT_RECOVERY_AGE — hours before item is eligible for recovery (default: 6)
|
||||
# SONARR_VERSION_MAJOR — expected Sonarr major version (e.g. 4)
|
||||
# RADARR_VERSION_MAJOR — expected Radarr major version (e.g. 6)
|
||||
# LIDARR_VERSION_MAJOR — expected Lidarr major version (e.g. 3)
|
||||
# ARR_RECOVERY_STATS — stats file path (read by coffee report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arrs_failed_stalled_recovery.sh — normal run
|
||||
# arrs_failed_stalled_recovery.sh --dry-run — show what would be actioned, no changes
|
||||
# arrs_failed_stalled_recovery.sh --log — verbose output
|
||||
# arrs_failed_stalled_recovery.sh --status — show config and exit
|
||||
#
|
||||
# Recommended schedule: 0 5 * * * (5am daily)
|
||||
# Or every 6hr: 0 */6 * * * (matches ARR_IMPORT_RECOVERY_AGE default)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases SONARR_*, RADARR_*, LIDARR_* vars
|
||||
detect_hosts
|
||||
|
||||
# jq is required — not optional — for JSON parsing
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq is not installed — required for arr API JSON parsing"
|
||||
error "Install: apt-get install jq or brew install jq"
|
||||
notify "arrs_failed_stalled_recovery failed on $(hostname) — jq not installed" \
|
||||
"Arr Recovery" "warning"
|
||||
exit 1
|
||||
fi
|
||||
log "jq found"
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no items will be blocklisted or searched"
|
||||
|
||||
# Age threshold in seconds
|
||||
AGE_THRESHOLD_SECONDS=$(( ARR_IMPORT_RECOVERY_AGE * 3600 ))
|
||||
|
||||
TOTAL_ACTIONED=0
|
||||
TOTAL_SKIPPED=0
|
||||
ARR_SUMMARIES=()
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Sonarr: ${SONARR_URL:-not configured} (recovery: ${SONARR_RECOVERY:-true})"
|
||||
echo "$ICON_SYNC Radarr: ${RADARR_URL:-not configured} (recovery: ${RADARR_RECOVERY:-true})"
|
||||
echo "$ICON_SYNC Lidarr: ${LIDARR_URL:-not configured on this host} (recovery: ${LIDARR_RECOVERY:-false})"
|
||||
echo "$ICON_TIME Age thresh: ${ARR_IMPORT_RECOVERY_AGE}hr"
|
||||
echo "$ICON_GEAR Sonarr ver: v${SONARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Radarr ver: v${RADARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Lidarr ver: v${LIDARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_NOTIFY Notify: unRAID=${NOTIFY_UNRAID:-false} Discord=$([[ -n "${MY_DISCORD_WEBHOOK:-}" ]] && echo enabled || echo disabled)"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
# Check if a queue item is older than ARR_IMPORT_RECOVERY_AGE
|
||||
# Returns 0 (old enough) or 1 (too new — skip)
|
||||
item_is_old_enough() {
|
||||
local added="$1"
|
||||
[[ -z "$added" ]] && return 0 # no date = treat as old enough, safe to act
|
||||
local added_epoch
|
||||
added_epoch=$(date -d "$added" +%s 2>/dev/null) || return 0
|
||||
local age_seconds=$(( $(date +%s) - added_epoch ))
|
||||
[[ "$age_seconds" -ge "$AGE_THRESHOLD_SECONDS" ]]
|
||||
}
|
||||
|
||||
# Query the arr queue API and return all records
|
||||
# Args: url, api_key, api_version
|
||||
get_queue_data() {
|
||||
local url="$1" api_key="$2" api_version="$3"
|
||||
curl -sf --max-time 15 \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/queue?page=1&pageSize=200&includeUnknownSeriesItems=true&includeUnknownArtistItems=true" \
|
||||
2>/dev/null
|
||||
}
|
||||
|
||||
# Blocklist and remove a queue item
|
||||
# Args: url, api_key, api_version, queue_id
|
||||
blocklist_item() {
|
||||
local url="$1" api_key="$2" api_version="$3" queue_id="$4"
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would blocklist queue item $queue_id"
|
||||
return 0
|
||||
fi
|
||||
curl -sf --max-time 15 \
|
||||
-X DELETE \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
"${url}/api/${api_version}/queue/${queue_id}?removeFromClient=true&blocklist=true&skipRedownload=false" \
|
||||
>/dev/null 2>&1
|
||||
}
|
||||
|
||||
# Trigger a new search for the media item
|
||||
# Args: url, api_key, api_version, arr_type, media_id
|
||||
trigger_search() {
|
||||
local url="$1" api_key="$2" api_version="$3" arr_type="$4" media_id="$5"
|
||||
local command body
|
||||
case "$arr_type" in
|
||||
sonarr) command="EpisodeSearch"; body="{\"name\":\"EpisodeSearch\",\"episodeIds\":[$media_id]}" ;;
|
||||
radarr) command="MoviesSearch"; body="{\"name\":\"MoviesSearch\",\"movieIds\":[$media_id]}" ;;
|
||||
lidarr) command="AlbumSearch"; body="{\"name\":\"AlbumSearch\",\"albumIds\":[$media_id]}" ;;
|
||||
esac
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would trigger $command for media ID $media_id"
|
||||
return 0
|
||||
fi
|
||||
curl -sf --max-time 15 \
|
||||
-X POST \
|
||||
-H "X-Api-Key: $api_key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body" \
|
||||
"${url}/api/${api_version}/command" \
|
||||
>/dev/null 2>&1
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ── PROCESS AN ARR ────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
# Args: display_name, arr_type, url, api_key, api_version, enabled,
|
||||
# version_major, version_api_prefix
|
||||
#
|
||||
# Exits cleanly if disabled.
|
||||
# Checks API reachability and version before touching queue.
|
||||
# Processes each problem item: blocklist + trigger new search.
|
||||
# Silent when clean — only warns when problems found or actioned.
|
||||
|
||||
process_arr() {
|
||||
local arr_name="$1"
|
||||
local arr_type="$2"
|
||||
local url="$3"
|
||||
local api_key="$4"
|
||||
local api_version="$5"
|
||||
local enabled="$6"
|
||||
local version_major="$7"
|
||||
local version_api_prefix="$8"
|
||||
|
||||
local actioned=0 skipped_new=0
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC $arr_name ━━━"
|
||||
|
||||
# Disabled — skip cleanly
|
||||
if [[ "$enabled" != "true" ]]; then
|
||||
log "$arr_name recovery disabled — skipping"
|
||||
ARR_SUMMARIES+=("$arr_name: disabled")
|
||||
return
|
||||
fi
|
||||
|
||||
# URL not configured on this host — skip cleanly
|
||||
if [[ -z "$url" ]]; then
|
||||
log "$arr_name not configured on $MY_ID — skipping"
|
||||
ARR_SUMMARIES+=("$arr_name: not configured on $MY_ID")
|
||||
return
|
||||
fi
|
||||
|
||||
# API reachability
|
||||
if ! check_api "$url" "$arr_name" 10; then
|
||||
warn "$arr_name API unreachable — skipping"
|
||||
ARR_SUMMARIES+=("$arr_name: unreachable")
|
||||
return
|
||||
fi
|
||||
|
||||
# Version check — exit if API structure may have changed
|
||||
if ! check_arr_version "$url" "$api_key" "$version_api_prefix" \
|
||||
"$version_major" "$arr_name"; then
|
||||
ARR_SUMMARIES+=("$arr_name: version mismatch — skipped")
|
||||
return
|
||||
fi
|
||||
|
||||
# Fetch queue
|
||||
local queue_data
|
||||
queue_data=$(get_queue_data "$url" "$api_key" "$api_version")
|
||||
if [[ -z "$queue_data" ]]; then
|
||||
warn "$arr_name — could not retrieve queue data"
|
||||
ARR_SUMMARIES+=("$arr_name: queue fetch failed")
|
||||
return
|
||||
fi
|
||||
|
||||
local total_records
|
||||
total_records=$(echo "$queue_data" | jq '.totalRecords // 0' 2>/dev/null)
|
||||
log "$arr_name queue: $total_records total items"
|
||||
|
||||
# Filter for problem items — never touch downloading or imported
|
||||
local problem_items
|
||||
problem_items=$(echo "$queue_data" | jq -c '
|
||||
.records // [] |
|
||||
.[] |
|
||||
select(
|
||||
.trackedDownloadState != "downloading" and
|
||||
.trackedDownloadState != "imported" and
|
||||
(
|
||||
.trackedDownloadState == "importFailed" or
|
||||
.trackedDownloadState == "importPending" or
|
||||
.trackedDownloadStatus == "error" or
|
||||
(.status == "warning" and (
|
||||
(.errorMessage // "" | ascii_downcase | contains("stalled")) or
|
||||
(.statusMessages // [] | .[] | .messages // [] | .[] |
|
||||
ascii_downcase | contains("stalled"))
|
||||
))
|
||||
)
|
||||
)
|
||||
' 2>/dev/null)
|
||||
|
||||
if [[ -z "$problem_items" ]]; then
|
||||
log "$arr_name — clean ✅ no failed imports or stalled downloads"
|
||||
ARR_SUMMARIES+=("$arr_name: clean ✅")
|
||||
return
|
||||
fi
|
||||
|
||||
local problem_count
|
||||
problem_count=$(echo "$problem_items" | wc -l)
|
||||
warn "$arr_name — found $problem_count problem item(s)"
|
||||
|
||||
# Process each problem item
|
||||
while IFS= read -r item; do
|
||||
[[ -z "$item" ]] && continue
|
||||
|
||||
local queue_id title added tracked_state tracked_status problem_type media_id
|
||||
|
||||
queue_id=$(echo "$item" | jq -r '.id // empty' 2>/dev/null)
|
||||
title=$(echo "$item" | jq -r '.title // "Unknown"' 2>/dev/null)
|
||||
added=$(echo "$item" | jq -r '.added // empty' 2>/dev/null)
|
||||
tracked_state=$(echo "$item" | jq -r '.trackedDownloadState // ""' 2>/dev/null)
|
||||
tracked_status=$(echo "$item" | jq -r '.trackedDownloadStatus // ""' 2>/dev/null)
|
||||
|
||||
# Human-readable problem type
|
||||
case "$tracked_state" in
|
||||
importFailed) problem_type="import failed" ;;
|
||||
importPending) problem_type="import pending/stuck" ;;
|
||||
*)
|
||||
[[ "$tracked_status" == "error" ]] && \
|
||||
problem_type="error" || problem_type="stalled"
|
||||
;;
|
||||
esac
|
||||
|
||||
# Media ID for search trigger
|
||||
case "$arr_type" in
|
||||
sonarr) media_id=$(echo "$item" | jq -r '.episodeId // .episode.id // empty' 2>/dev/null) ;;
|
||||
radarr) media_id=$(echo "$item" | jq -r '.movieId // .movie.id // empty' 2>/dev/null) ;;
|
||||
lidarr) media_id=$(echo "$item" | jq -r '.albumId // .album.id // empty' 2>/dev/null) ;;
|
||||
esac
|
||||
|
||||
[[ -z "$queue_id" ]] && continue
|
||||
|
||||
# Age check — skip items that are too new to have self-resolved
|
||||
if ! item_is_old_enough "$added"; then
|
||||
log " Skipping (too new < ${ARR_IMPORT_RECOVERY_AGE}hr): $title"
|
||||
(( skipped_new++ ))
|
||||
(( TOTAL_SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
warn " $ICON_TRASH $problem_type — $title"
|
||||
|
||||
# Step 1: Blocklist + remove from queue
|
||||
if ! blocklist_item "$url" "$api_key" "$api_version" "$queue_id"; then
|
||||
warn " Failed to blocklist: $title"
|
||||
(( TOTAL_SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
log " Blocklisted: $queue_id"
|
||||
|
||||
# Step 2: Trigger new search
|
||||
if [[ -n "$media_id" ]]; then
|
||||
if trigger_search "$url" "$api_key" "$api_version" "$arr_type" "$media_id"; then
|
||||
log " New search triggered: $title"
|
||||
else
|
||||
warn " Blocklisted but search trigger failed: $title"
|
||||
fi
|
||||
else
|
||||
warn " Blocklisted but no media ID found — search not triggered: $title"
|
||||
fi
|
||||
|
||||
(( actioned++ ))
|
||||
(( TOTAL_ACTIONED++ ))
|
||||
|
||||
done <<< "$problem_items"
|
||||
|
||||
if [[ "$actioned" -gt 0 ]]; then
|
||||
warn "$arr_name — actioned: $actioned | skipped (too new): $skipped_new"
|
||||
else
|
||||
log "$arr_name — nothing actioned | skipped (too new): $skipped_new"
|
||||
fi
|
||||
|
||||
ARR_SUMMARIES+=("$arr_name: actioned $actioned | too new $skipped_new")
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Process Each Arr ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Arrs Failed/Stalled Recovery — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
echo "$ICON_HOST $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
log "Age threshold: ${ARR_IMPORT_RECOVERY_AGE}hr"
|
||||
|
||||
START=$(date +%s)
|
||||
|
||||
# Sonarr — uses aliased vars set by detect_hosts()
|
||||
process_arr \
|
||||
"Sonarr" \
|
||||
"sonarr" \
|
||||
"${SONARR_URL:-}" \
|
||||
"${SONARR_API_KEY:-}" \
|
||||
"v3" \
|
||||
"${SONARR_RECOVERY:-true}" \
|
||||
"${SONARR_VERSION_MAJOR:-4}" \
|
||||
"v3"
|
||||
|
||||
# Radarr — uses aliased vars set by detect_hosts()
|
||||
process_arr \
|
||||
"Radarr" \
|
||||
"radarr" \
|
||||
"${RADARR_URL:-}" \
|
||||
"${RADARR_API_KEY:-}" \
|
||||
"v3" \
|
||||
"${RADARR_RECOVERY:-true}" \
|
||||
"${RADARR_VERSION_MAJOR:-6}" \
|
||||
"v3"
|
||||
|
||||
# Lidarr — HOST1 only, LIDARR_URL empty on HOST2 → exits cleanly via "not configured" guard
|
||||
process_arr \
|
||||
"Lidarr" \
|
||||
"lidarr" \
|
||||
"${LIDARR_URL:-}" \
|
||||
"${LIDARR_API_KEY:-}" \
|
||||
"v1" \
|
||||
"${LIDARR_RECOVERY:-false}" \
|
||||
"${LIDARR_VERSION_MAJOR:-3}" \
|
||||
"v1"
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY ARR RECOVERY SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo "$ICON_TRASH Actioned: $TOTAL_ACTIONED items blocklisted + searched"
|
||||
echo "$ICON_SKIP Skipped: $TOTAL_SKIPPED items (too new)"
|
||||
echo ""
|
||||
for summary in "${ARR_SUMMARIES[@]}"; do
|
||||
echo " $ICON_SUMMARY $summary"
|
||||
done
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ "$TOTAL_ACTIONED" -gt 0 ]]; then
|
||||
warn "$ICON_DONE Done — $TOTAL_ACTIONED item(s) blocklisted and re-searched"
|
||||
notify "Arr recovery on $(hostname) — $TOTAL_ACTIONED item(s) blocklisted and re-searched" \
|
||||
"Arr Recovery" "warning"
|
||||
else
|
||||
echo "$ICON_DONE Done — nothing to recover (all arrs clean)"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# Write stats for sunday_morning_coffee_report.sh
|
||||
if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_RECOVERY_STATS:-}" ]]; then
|
||||
echo "$(date '+%Y-%m-%d')|$(date '+%H:%M')|${TOTAL_ACTIONED}|${TOTAL_SKIPPED}" \
|
||||
>> "$ARR_RECOVERY_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
@@ -1,452 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Lidarr Missing Art =========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Fetch missing album and artist artwork for the Lidarr music library. Downloads
|
||||
# only what is absent — never overwrites existing files. Idempotent re-runs are
|
||||
# safe.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Artwork targets per album folder: cover.jpg cdart.png back.jpg
|
||||
# Artwork targets per artist folder: folder.jpg fanart.jpg logo.png banner.jpg
|
||||
#
|
||||
# Sources (tried in order, first success wins):
|
||||
# Album covers: fanart.tv → iTunes fallback
|
||||
# Artist art: fanart.tv → Deezer fallback → Last.fm fallback
|
||||
#
|
||||
# Reads from Lidarr API only — no writes back to Lidarr. Never modifies audio
|
||||
# tags or renames media files. Only writes missing artwork files to existing
|
||||
# album/artist directories.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# acquire_lock — prevents concurrent runs during large library scans
|
||||
# curl + jq check — fail fast if tools missing
|
||||
# API reachability — verified before processing begins
|
||||
# detect_hosts() — exits cleanly if LIDARR_URL empty (HOST2, no Lidarr)
|
||||
# Skip existing — never overwrites, idempotent re-runs are safe
|
||||
# Min file size check — rejects corrupt/placeholder downloads (LIDARR_ART_MIN_SIZE)
|
||||
# Parallel job cap — LIDARR_ART_MAX_PARALLEL — avoids hammering external APIs
|
||||
# Download retries — LIDARR_ART_RETRIES attempts per image before giving up
|
||||
# Rate limiting — LIDARR_ART_SLEEP_BETWEEN between fanart.tv API calls
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST1_LIDARR_URL / HOST1_LIDARR_API_KEY
|
||||
# Aliased by detect_hosts() — script uses LIDARR_URL / LIDARR_API_KEY
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# FANART_API_KEY — fanart.tv API key
|
||||
# LASTFM_API_KEY — last.fm API key
|
||||
# LIDARR_ART_MIN_SIZE — minimum valid download size in bytes
|
||||
# LIDARR_ART_MAX_PARALLEL — concurrent background download jobs
|
||||
# LIDARR_ART_RETRIES — download retry attempts per image
|
||||
# LIDARR_ART_SLEEP_BETWEEN — seconds between fanart.tv API calls
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# lidarr_missing_art.sh — fetch all missing artwork
|
||||
# lidarr_missing_art.sh --dry-run — preview without downloading
|
||||
# lidarr_missing_art.sh --log — verbose per-item output
|
||||
# lidarr_missing_art.sh --status — show config and exit
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for API calls"
|
||||
exit 1
|
||||
fi
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# Build path map for translate_path() — container path → host path
|
||||
declare -A ARR_PATH_MAP
|
||||
for _key in "${!HOST1_LIDARR_PATH_MAP[@]}"; do
|
||||
ARR_PATH_MAP["$_key"]="${HOST1_LIDARR_PATH_MAP[$_key]}"
|
||||
done
|
||||
unset _key
|
||||
|
||||
# HOST guard — Lidarr runs on HOST1 only
|
||||
if [[ -z "$LIDARR_URL" ]]; then
|
||||
log "Lidarr not configured for $MY_ID — nothing to do"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
info "$MY_ID ($LOCAL_SERVER_NAME) — tools OK"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Lidarr URL: $LIDARR_URL"
|
||||
echo "$ICON_NET Fanart key: $([[ -n "${FANART_API_KEY:-}" ]] && echo "set" || echo "not set")"
|
||||
echo "$ICON_NET LastFM key: $([[ -n "${LASTFM_API_KEY:-}" ]] && echo "set" || echo "not set")"
|
||||
echo "$ICON_GEAR Min size: ${LIDARR_ART_MIN_SIZE} bytes"
|
||||
echo "$ICON_GEAR Parallel: $LIDARR_ART_MAX_PARALLEL jobs"
|
||||
echo "$ICON_RETRY Retries: $LIDARR_ART_RETRIES"
|
||||
echo "$ICON_TIME API sleep: ${LIDARR_ART_SLEEP_BETWEEN}s"
|
||||
echo "$ICON_NOTIFY Notify: unRAID=${NOTIFY_UNRAID:-false} Discord=$([[ -n "${MY_DISCORD_WEBHOOK:-}" ]] && echo enabled || echo disabled)"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be written"
|
||||
|
||||
# ── Temp dir for subshell fetch/fail counters ─────────────────────────────────────────────────
|
||||
LIDARR_TMP=$(mktemp -d)
|
||||
trap 'rm -rf "$LIDARR_TMP"' EXIT
|
||||
touch "$LIDARR_TMP/album_fetches" "$LIDARR_TMP/album_fails" \
|
||||
"$LIDARR_TMP/artist_fetches" "$LIDARR_TMP/artist_fails"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── FUNCTIONS ─────────────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
curl_json() {
|
||||
curl -s --connect-timeout 5 --max-time 20 "$1"
|
||||
}
|
||||
|
||||
job_count() {
|
||||
jobs -rp | wc -l
|
||||
}
|
||||
|
||||
wait_for_slot() {
|
||||
while (( $(job_count) >= LIDARR_ART_MAX_PARALLEL )); do
|
||||
sleep 0.2
|
||||
done
|
||||
}
|
||||
|
||||
# Downloads URL to dest only if dest doesn't exist and downloaded size >= MIN_SIZE.
|
||||
# Returns 0 on success or skip (file already exists), 1 on failure.
|
||||
download_if_valid() {
|
||||
local url="$1"
|
||||
local dest="$2"
|
||||
|
||||
[[ -z "$url" || "$url" == "null" ]] && return 1
|
||||
[[ -f "$dest" ]] && return 0
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
log "DRY RUN — would fetch: $(basename "$dest")"
|
||||
return 0
|
||||
fi
|
||||
|
||||
local tmp="${dest}.tmp"
|
||||
local i
|
||||
for (( i=0; i<=LIDARR_ART_RETRIES; i++ )); do
|
||||
curl -s --connect-timeout 5 --max-time 20 -L -o "$tmp" "$url"
|
||||
local size
|
||||
size=$(stat -c%s "$tmp" 2>/dev/null || echo 0)
|
||||
if (( size > LIDARR_ART_MIN_SIZE )); then
|
||||
mv "$tmp" "$dest"
|
||||
log " Fetched: $(basename "$dest")"
|
||||
return 0
|
||||
fi
|
||||
rm -f "$tmp"
|
||||
sleep 1
|
||||
done
|
||||
|
||||
warn "Failed to fetch: $(basename "$dest")"
|
||||
return 1
|
||||
}
|
||||
|
||||
deezer_artist_image() {
|
||||
local artist="$1"
|
||||
local query
|
||||
query=$(printf "%s" "$artist" | sed 's/ /+/g')
|
||||
curl_json "https://api.deezer.com/search/artist?q=$query" |
|
||||
jq -r '.data[0].picture_xl // empty' 2>/dev/null
|
||||
}
|
||||
|
||||
lastfm_artist_image() {
|
||||
local artist="$1"
|
||||
local encoded
|
||||
encoded=$(printf "%s" "$artist" | sed 's/ /%20/g')
|
||||
curl_json "https://ws.audioscrobbler.com/2.0/?method=artist.getinfo&artist=$encoded&api_key=$LASTFM_API_KEY&format=json" |
|
||||
jq -r '.artist.image[-1]["#text"] // empty' 2>/dev/null
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Verify Lidarr reachable ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_NET Lidarr API ━━━"
|
||||
|
||||
if ! curl_json "$LIDARR_URL/api/v1/system/status?apikey=$LIDARR_API_KEY" | jq -e '.version' >/dev/null 2>&1; then
|
||||
error "Lidarr API unreachable at $LIDARR_URL — aborting"
|
||||
notify "lidarr_missing_art failed — Lidarr API unreachable on $(hostname)" "Lidarr Missing Art" "warning"
|
||||
exit 1
|
||||
fi
|
||||
info "Lidarr reachable — $LIDARR_URL"
|
||||
|
||||
START=$(date +%s)
|
||||
|
||||
ALBUMS_CHECKED=0
|
||||
ALBUMS_COMPLETE=0
|
||||
|
||||
ARTISTS_CHECKED=0
|
||||
ARTISTS_COMPLETE=0
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Build Album Directory Map ━━━
|
||||
# ==============================================================================================
|
||||
# Lidarr's album API never populates .path — derive album dirs from track file paths instead.
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Building Album Directory Map ━━━"
|
||||
|
||||
declare -A ALBUM_DIR_MAP
|
||||
_artist_list=$(curl_json "$LIDARR_URL/api/v1/artist?apikey=$LIDARR_API_KEY")
|
||||
_map_artist_count=$(echo "$_artist_list" | jq '. | length')
|
||||
info "Fetching track files for $_map_artist_count artists..."
|
||||
|
||||
while IFS= read -r _artist_id; do
|
||||
[[ -z "$_artist_id" ]] && continue
|
||||
while IFS=$'\t' read -r _album_id _track_path; do
|
||||
[[ -z "$_album_id" || -z "$_track_path" || "$_track_path" == "null" ]] && continue
|
||||
ALBUM_DIR_MAP["$_album_id"]=$(dirname "$_track_path")
|
||||
done < <(curl_json "$LIDARR_URL/api/v1/trackFile?artistId=${_artist_id}&apikey=$LIDARR_API_KEY" | \
|
||||
jq -r '.[] | [(.albumId | tostring), .path] | @tsv' 2>/dev/null)
|
||||
done < <(echo "$_artist_list" | jq -r '.[].id')
|
||||
unset _artist_list _map_artist_count _artist_id _album_id _track_path
|
||||
|
||||
info "Mapped ${#ALBUM_DIR_MAP[@]} albums with local tracks"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Albums ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_EMBY Albums ━━━"
|
||||
|
||||
albums=$(curl_json "$LIDARR_URL/api/v1/album?apikey=$LIDARR_API_KEY")
|
||||
|
||||
if [[ -z "$albums" || "$albums" == "null" ]]; then
|
||||
error "Lidarr album API returned empty — aborting"
|
||||
notify "lidarr_missing_art failed — album API empty on $(hostname)" "Lidarr Missing Art" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
total_albums=$(echo "$albums" | jq '. | length')
|
||||
info "Processing $total_albums albums..."
|
||||
|
||||
while IFS=$'\t' read -r mbid artist_name album_name album_id; do
|
||||
raw_dir="${ALBUM_DIR_MAP[$album_id]:-}"
|
||||
[[ -z "$raw_dir" ]] && continue # not downloaded, skip
|
||||
local_path=$(translate_path "$raw_dir")
|
||||
(( ALBUMS_CHECKED++ ))
|
||||
|
||||
[[ ! -d "$local_path" ]] && continue
|
||||
|
||||
log "[$ALBUMS_CHECKED/$total_albums] $artist_name — $album_name"
|
||||
|
||||
if [[ -f "$local_path/cover.jpg" &&
|
||||
-f "$local_path/cdart.png" &&
|
||||
-f "$local_path/back.jpg" ]]; then
|
||||
(( ALBUMS_COMPLETE++ ))
|
||||
log " complete — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
wait_for_slot
|
||||
|
||||
(
|
||||
_fetches=0 _fails=0
|
||||
|
||||
JSON=""
|
||||
if [[ -n "$mbid" && "$mbid" != "null" ]]; then
|
||||
JSON=$(curl_json "http://webservice.fanart.tv/v3/music/albums/$mbid?api_key=$FANART_API_KEY")
|
||||
sleep "$LIDARR_ART_SLEEP_BETWEEN"
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/cover.jpg" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.[].albumcover[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/cover.jpg"; then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
query=$(printf "%s %s" "$artist_name" "$album_name" | sed 's/ /+/g')
|
||||
itunes=$(curl_json "https://itunes.apple.com/search?term=$query&entity=album&limit=1" |
|
||||
jq -r '.results[0].artworkUrl100 // empty' 2>/dev/null | sed 's/100x100/600x600/')
|
||||
if download_if_valid "$itunes" "$local_path/cover.jpg"; then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
(( _fails++ ))
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/cdart.png" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.[].cdart[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/cdart.png"; then (( _fetches++ )); else (( _fails++ )); fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/back.jpg" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.[].albumback[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/back.jpg"; then (( _fetches++ )); else (( _fails++ )); fi
|
||||
fi
|
||||
|
||||
(( _fetches > 0 )) && printf '1\n' >> "$LIDARR_TMP/album_fetches"
|
||||
(( _fails > 0 )) && printf '1\n' >> "$LIDARR_TMP/album_fails"
|
||||
) &
|
||||
|
||||
done < <(echo "$albums" | jq -r '.[] | [(.foreignAlbumId // ""), (.artist.artistName // ""), (.title // ""), (.id | tostring)] | @tsv')
|
||||
|
||||
wait
|
||||
|
||||
ALBUM_FETCHED=$(wc -l < "$LIDARR_TMP/album_fetches" 2>/dev/null || echo 0)
|
||||
ALBUM_FAILED=$(wc -l < "$LIDARR_TMP/album_fails" 2>/dev/null || echo 0)
|
||||
ALBUM_MISSING=$(( ALBUMS_CHECKED - ALBUMS_COMPLETE ))
|
||||
info "Checked: $ALBUMS_CHECKED | Complete: $ALBUMS_COMPLETE | Needed art: $ALBUM_MISSING | Fetched: $ALBUM_FETCHED | Failed: $ALBUM_FAILED"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Artists ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_EMBY Artists ━━━"
|
||||
|
||||
artists=$(curl_json "$LIDARR_URL/api/v1/artist?apikey=$LIDARR_API_KEY")
|
||||
|
||||
if [[ -z "$artists" || "$artists" == "null" ]]; then
|
||||
error "Lidarr artist API returned empty — aborting"
|
||||
notify "lidarr_missing_art failed — artist API empty on $(hostname)" "Lidarr Missing Art" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
total_artists=$(echo "$artists" | jq '. | length')
|
||||
info "Processing $total_artists artists..."
|
||||
|
||||
while IFS=$'\t' read -r local_path mbid name; do
|
||||
local_path=$(translate_path "$local_path")
|
||||
(( ARTISTS_CHECKED++ ))
|
||||
|
||||
[[ ! -d "$local_path" ]] && continue
|
||||
[[ -z "$mbid" || "$mbid" == "null" ]] && continue
|
||||
|
||||
log "[$ARTISTS_CHECKED/$total_artists] $name"
|
||||
|
||||
if [[ -f "$local_path/folder.jpg" &&
|
||||
-f "$local_path/fanart.jpg" &&
|
||||
-f "$local_path/logo.png" &&
|
||||
-f "$local_path/banner.jpg" ]]; then
|
||||
(( ARTISTS_COMPLETE++ ))
|
||||
log " complete — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
wait_for_slot
|
||||
|
||||
(
|
||||
_fetches=0 _fails=0
|
||||
|
||||
JSON=$(curl_json "http://webservice.fanart.tv/v3/music/$mbid?api_key=$FANART_API_KEY")
|
||||
sleep "$LIDARR_ART_SLEEP_BETWEEN"
|
||||
|
||||
if [[ ! -f "$local_path/folder.jpg" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.artistthumb[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/folder.jpg"; then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
IMG=$(deezer_artist_image "$name")
|
||||
if download_if_valid "$IMG" "$local_path/folder.jpg"; then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
IMG=$(lastfm_artist_image "$name")
|
||||
if download_if_valid "$IMG" "$local_path/folder.jpg"; then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
(( _fails++ ))
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/fanart.jpg" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.artistbackground[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/fanart.jpg"; then
|
||||
(( _fetches++ ))
|
||||
else
|
||||
IMG=$(deezer_artist_image "$name")
|
||||
if download_if_valid "$IMG" "$local_path/fanart.jpg"; then (( _fetches++ )); else (( _fails++ )); fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/logo.png" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.hdmusiclogo[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/logo.png"; then (( _fetches++ )); else (( _fails++ )); fi
|
||||
fi
|
||||
|
||||
if [[ ! -f "$local_path/banner.jpg" ]]; then
|
||||
IMG=$(echo "$JSON" | jq -r '.musicbanner[0].url // empty' 2>/dev/null)
|
||||
if download_if_valid "$IMG" "$local_path/banner.jpg"; then (( _fetches++ )); else (( _fails++ )); fi
|
||||
fi
|
||||
|
||||
(( _fetches > 0 )) && printf '1\n' >> "$LIDARR_TMP/artist_fetches"
|
||||
(( _fails > 0 )) && printf '1\n' >> "$LIDARR_TMP/artist_fails"
|
||||
) &
|
||||
|
||||
done < <(echo "$artists" | jq -r '.[] | [.path, .foreignArtistId, .artistName] | @tsv')
|
||||
|
||||
wait
|
||||
|
||||
ARTIST_FETCHED=$(wc -l < "$LIDARR_TMP/artist_fetches" 2>/dev/null || echo 0)
|
||||
ARTIST_FAILED=$(wc -l < "$LIDARR_TMP/artist_fails" 2>/dev/null || echo 0)
|
||||
ARTIST_MISSING=$(( ARTISTS_CHECKED - ARTISTS_COMPLETE ))
|
||||
info "Checked: $ARTISTS_CHECKED | Complete: $ARTISTS_COMPLETE | Needed art: $ARTIST_MISSING | Fetched: $ARTIST_FETCHED | Failed: $ARTIST_FAILED"
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY LIDARR MISSING ART SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $((END - START)))"
|
||||
echo "$ICON_EMBY Albums: $ALBUMS_CHECKED checked | $ALBUMS_COMPLETE complete | $ALBUM_FETCHED fetched | $ALBUM_FAILED failed"
|
||||
echo "$ICON_EMBY Artists: $ARTISTS_CHECKED checked | $ARTISTS_COMPLETE complete | $ARTIST_FETCHED fetched | $ARTIST_FAILED failed"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo "$ICON_WARN Status: DRY RUN — no files written"
|
||||
else
|
||||
echo "$ICON_DONE Status: $ICON_SUCCESS ALL DONE"
|
||||
notify "Lidarr art fetch complete on $(hostname) — ${ALBUMS_CHECKED} albums, ${ARTISTS_CHECKED} artists processed" "Lidarr Missing Art" "normal"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
exit 0
|
||||
Regular → Executable
+72
-11
@@ -43,14 +43,66 @@
|
||||
# Media profile also removes: *.iso *.lrc
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Pre-Scan Cleanup
|
||||
# Runs before arr cleanup scripts so orphan detection only encounters actual
|
||||
# media files. Scene debris and download artifacts would otherwise appear as
|
||||
# untracked files and inflate false-positive orphan counts.
|
||||
#
|
||||
# Profile Separation
|
||||
# Anime and media share different cleanup patterns because their content
|
||||
# differs. *.lrc (lyrics) and *.iso belong in media cleanup but not anime.
|
||||
# Separate profiles prevent cross-contamination of rules.
|
||||
#
|
||||
# Conservative by Default
|
||||
# Only explicitly listed patterns are removed. The script never guesses
|
||||
# at file intent — if a pattern is not in the list, the file is untouched.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# acquire_lock "wait" — wait if previous run still active
|
||||
# detect_hosts() — correct folder lists per host via MY_ID aliases
|
||||
# Empty array guards — warns and exits cleanly if no folders or patterns configured
|
||||
# Folder existence — skips missing folders with warning, continues others
|
||||
# validate_unraid_cmd — notify script validated before use
|
||||
# Root Enforcement
|
||||
# Media files are owned by container users; deleting them requires root.
|
||||
#
|
||||
# Profile Required
|
||||
# Exits with usage if no profile is given. There is no default profile — an
|
||||
# unspecified profile must never fall through to cleaning something.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock "wait" — waits for a previous run to finish rather than
|
||||
# skipping, so a long anime pass does not cause the media pass to be dropped.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases HOST*_ANIME_CLEAN_FOLDERS / HOST*_MEDIA_CLEAN_FOLDERS
|
||||
# to the correct host's values.
|
||||
#
|
||||
# Empty Array Guards
|
||||
# Exits cleanly if the resolved folder list or pattern list is empty. An empty
|
||||
# pattern list would otherwise build a find with no -iname terms and match
|
||||
# every file in the tree.
|
||||
#
|
||||
# Clean Path Depth Guard
|
||||
# Every folder must be an absolute path at least three levels deep before it is
|
||||
# scanned. The patterns include *.sh, *.zip, *.rar and *.exe, so a truncated
|
||||
# entry like /mnt/user — which passes an existence check — would delete
|
||||
# matching files across every share on the array.
|
||||
#
|
||||
# Folder Existence
|
||||
# Missing folders are skipped with a warning; remaining folders still process.
|
||||
#
|
||||
# Explicit Pattern List
|
||||
# Only patterns named in ANIME_FILE_PATTERNS / MEDIA_FILE_PATTERNS are removed.
|
||||
# The script never infers intent from file size, age, or location.
|
||||
#
|
||||
# Count Before Delete
|
||||
# Matching files are counted first; a folder with zero matches short-circuits
|
||||
# before any rm is constructed.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run lists every file that would be deleted and removes nothing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -117,10 +169,6 @@ acquire_lock "wait"
|
||||
# detect_hosts() sets MY_ID and aliases HOST*_ANIME/MEDIA_CLEAN_FOLDERS
|
||||
detect_hosts
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
# Resolve profile folders and patterns
|
||||
case "$PROFILE" in
|
||||
@@ -191,6 +239,19 @@ for FOLDER in "${CLEAN_FOLDERS[@]}"; do
|
||||
FOLDER_NAME=$(basename "$FOLDER")
|
||||
echo "━━━ $ICON_CLEAN $FOLDER_NAME ━━━"
|
||||
|
||||
# The pattern list includes *.sh, *.zip, *.rar and *.exe. A truncated entry such as
|
||||
# /mnt/user passes the -d check below and would sweep every share on the array, so
|
||||
# require an absolute path at least three levels deep before scanning anything.
|
||||
_depth="${FOLDER//[^\/]/}"
|
||||
if [[ -z "$FOLDER" || "$FOLDER" != /* || "${#_depth}" -lt 3 ]]; then
|
||||
error "Refusing to clean unsafe path: '${FOLDER:-empty}' — expected an absolute path at least 3 levels deep"
|
||||
notify "Media cleaner ($PROFILE) refused unsafe path on $(hostname): '${FOLDER:-empty}'" \
|
||||
"Media Cleaner" "warning"
|
||||
FAILED+=("${FOLDER_NAME:-empty}")
|
||||
echo ""
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ ! -d "$FOLDER" ]]; then
|
||||
warn "$FOLDER_NAME not found — skipping"
|
||||
SKIPPED+=("$FOLDER_NAME")
|
||||
@@ -212,7 +273,7 @@ for FOLDER in "${CLEAN_FOLDERS[@]}"; do
|
||||
FILE_COUNT=$("${CMD[@]}" 2>/dev/null | wc -l)
|
||||
|
||||
if [[ "$FILE_COUNT" -eq 0 ]]; then
|
||||
log "$FOLDER_NAME — clean ✅"
|
||||
echo "$FOLDER_NAME — clean ✅"
|
||||
echo ""
|
||||
continue
|
||||
fi
|
||||
@@ -227,7 +288,7 @@ for FOLDER in "${CLEAN_FOLDERS[@]}"; do
|
||||
else
|
||||
CLEAN_CMD=("${CMD[@]}" -exec rm -f {} +)
|
||||
if "${CLEAN_CMD[@]}" 2>/dev/null; then
|
||||
log "$FOLDER_NAME — $FILE_COUNT file(s) removed"
|
||||
echo "$FOLDER_NAME — $FILE_COUNT file(s) removed"
|
||||
TOTAL_REMOVED=$(( TOTAL_REMOVED + FILE_COUNT ))
|
||||
else
|
||||
error "$FOLDER_NAME — cleanup failed"
|
||||
|
||||
Regular → Executable
+127
-18
@@ -14,16 +14,98 @@
|
||||
# or unRAID environment resets after updates.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# For each share in MEDIA_PERMISSION_SHARES:
|
||||
#
|
||||
# 1. Path safety and existence
|
||||
# → unsafe or missing paths are refused or skipped, never scanned
|
||||
#
|
||||
# 2. Count wrong ownership (diagnostic)
|
||||
# → find ! -user / ! -group — the number reported as "corrected"
|
||||
#
|
||||
# 3. Ownership pass — only if the count is non-zero
|
||||
# → chown PERMISSIONS_OWNER on non-matching entries only
|
||||
#
|
||||
# 4. Directory mode pass
|
||||
# → chmod PERMISSIONS_DIR_MODE on directories not already at that mode
|
||||
#
|
||||
# 5. File mode pass
|
||||
# → chmod PERMISSIONS_FILE_MODE on files not already at that mode
|
||||
# → "No such file" errors ignored: volatile dirs (Emby transcodes) race
|
||||
#
|
||||
# Every pass is conditional by design — see Conditional Passes below.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Daily Failsafe, Not Enforcer
|
||||
# Wrong ownership is a symptom of something else — a misconfigured container,
|
||||
# a manual copy, an rsync without --chown. This script corrects the symptom
|
||||
# daily rather than hunting the root cause. A persistent high correction count
|
||||
# is the signal to investigate the source.
|
||||
#
|
||||
# Runs First in the Window
|
||||
# Arr cleanup scripts depend on correct ownership to rename and delete files.
|
||||
# Permissions must be correct before cleanup runs — ordering is not optional.
|
||||
#
|
||||
# Separate Passes for Directories and Files
|
||||
# Directories need execute permission for traversal; files do not. Applying
|
||||
# the same mode to both is a common mistake this script avoids by design.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# acquire_lock "wait" — wait if previous run still active (large share scans)
|
||||
# detect_hosts() — correct share list per host via MY_ID aliases
|
||||
# Empty array guard — warns and exits cleanly if no shares configured
|
||||
# Folder existence — skips missing shares with warning, continues others
|
||||
# Separate passes — directories and files chmod'd separately for correctness
|
||||
# validate_unraid_cmd — notify script validated before use
|
||||
# Silent by default — only failures produce output, success is silent
|
||||
# Root Enforcement
|
||||
# chown to an arbitrary owner requires root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock "wait" — waits rather than skipping. Share scans are long, and
|
||||
# this runs first in the daily window; skipping it would let arr cleanup run
|
||||
# against uncorrected ownership.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases HOST*_MEDIA_PERMISSION_SHARES to this host's shares.
|
||||
#
|
||||
# Empty Array Guard
|
||||
# Exits cleanly if no shares are configured for this host.
|
||||
#
|
||||
# Share Path Depth Guard
|
||||
# Every share must be an absolute path at least three levels deep before it is
|
||||
# scanned. A truncated entry like /mnt/user passes an existence check and would
|
||||
# chown and chmod every share on the array — which, because chown/chmod restamp
|
||||
# ctime, would erase the age signal the arr cleanups depend on across the whole
|
||||
# library in a single run.
|
||||
#
|
||||
# Folder Existence
|
||||
# Missing shares are skipped with a warning; remaining shares still process.
|
||||
#
|
||||
# Separate Passes
|
||||
# Directories and files are chmod'd in separate passes — directories need the
|
||||
# execute bit for traversal, media files must not have it.
|
||||
#
|
||||
# Conditional Passes
|
||||
# Only entries whose owner or mode is actually wrong are touched. This is not
|
||||
# an optimisation: chown/chmod rewrite an inode's ctime even when the value is
|
||||
# unchanged, so a blanket pass would restamp every file nightly and destroy
|
||||
# ctime as an age signal. The arr cleanups gate orphan deletion on ctime, and
|
||||
# mtime cannot substitute — imports preserve the release's original timestamp.
|
||||
# Making any pass unconditional silently stops orphan collection.
|
||||
#
|
||||
# Transcode Race Tolerance
|
||||
# "No such file or directory" errors from the file pass are ignored. Volatile
|
||||
# directories such as Emby transcodes delete files mid-scan; that is expected,
|
||||
# not a permissions failure.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run counts the dirs, files and ownership entries that would change and
|
||||
# modifies nothing.
|
||||
#
|
||||
# Silent by Default
|
||||
# Only failures and diagnostics produce output; a clean run is quiet.
|
||||
#
|
||||
# Diagnostic — high corrected count on every run means a container has wrong PUID/PGID:
|
||||
# Correct values on unRAID: PUID=99 (nobody) PGID=100 (users)
|
||||
@@ -70,10 +152,6 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock "wait"
|
||||
|
||||
@@ -129,9 +207,25 @@ SKIPPED=()
|
||||
TOTAL_DIRS_FIXED=0
|
||||
TOTAL_FILES_FIXED=0
|
||||
|
||||
# Split for find's -user/-group predicates, which take them separately
|
||||
PERMISSIONS_USER="${PERMISSIONS_OWNER%%:*}"
|
||||
PERMISSIONS_GROUP="${PERMISSIONS_OWNER##*:}"
|
||||
|
||||
for SHARE in "${MEDIA_PERMISSION_SHARES[@]}"; do
|
||||
SHARE_NAME=$(basename "$SHARE")
|
||||
|
||||
# A truncated entry such as /mnt/user passes the -d check below and would chown/chmod
|
||||
# every share on the array. Because chown/chmod restamp ctime, that would erase the age
|
||||
# signal the arr cleanups gate orphan deletion on — across the whole library, in one run.
|
||||
_depth="${SHARE//[^\/]/}"
|
||||
if [[ -z "$SHARE" || "$SHARE" != /* || "${#_depth}" -lt 3 ]]; then
|
||||
error "Refusing to touch unsafe path: '${SHARE:-empty}' — expected an absolute path at least 3 levels deep"
|
||||
notify "Media permissions refused unsafe path on $(hostname): '${SHARE:-empty}'" \
|
||||
"Media Permissions" "warning"
|
||||
FAILED+=("${SHARE_NAME:-empty}")
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ ! -d "$SHARE" ]]; then
|
||||
warn "$SHARE_NAME not found — skipping"
|
||||
SKIPPED+=("$SHARE_NAME")
|
||||
@@ -144,7 +238,7 @@ for SHARE in "${MEDIA_PERMISSION_SHARES[@]}"; do
|
||||
2>/dev/null | wc -l)
|
||||
FILE_COUNT=$(find "$SHARE" -type f ! -perm "${PERMISSIONS_FILE_MODE:-664}" \
|
||||
2>/dev/null | wc -l)
|
||||
OWNER_COUNT=$(find "$SHARE" ! -user nobody -o ! -group users \
|
||||
OWNER_COUNT=$(find "$SHARE" \( ! -user "$PERMISSIONS_USER" -o ! -group "$PERMISSIONS_GROUP" \) \
|
||||
2>/dev/null | wc -l)
|
||||
warn "DRY RUN — $SHARE_NAME: $DIR_COUNT dirs, $FILE_COUNT files, $OWNER_COUNT ownership fixes needed"
|
||||
continue
|
||||
@@ -156,21 +250,36 @@ for SHARE in "${MEDIA_PERMISSION_SHARES[@]}"; do
|
||||
CHMOD_FILE_OK=true
|
||||
CHOWN_OK=true
|
||||
|
||||
# Every pass below is conditional — it touches only entries that are actually wrong.
|
||||
# This is not just an optimisation. chown/chmod rewrite an inode's ctime even when the
|
||||
# value is unchanged, so a blanket pass restamps every file in the share each night and
|
||||
# erases ctime as an age signal. The arr cleanups need that signal to tell a file that
|
||||
# just landed from one that has sat untracked for days — mtime can't do it, because an
|
||||
# import preserves the release's original timestamp (measured 2026-07-27: 400 of 400
|
||||
# files imported that week had mtimes over 7 days old, one of them 9613 days).
|
||||
|
||||
# Count files with wrong ownership before fixing (diagnostic)
|
||||
WRONG_OWNER=$(find "$SHARE" \( ! -user nobody -o ! -group users \) \
|
||||
WRONG_OWNER=$(find "$SHARE" \( ! -user "$PERMISSIONS_USER" -o ! -group "$PERMISSIONS_GROUP" \) \
|
||||
2>/dev/null | wc -l)
|
||||
|
||||
# Apply ownership first — affects all files and directories
|
||||
chown -R "$PERMISSIONS_OWNER" "$SHARE" 2>/dev/null || CHOWN_OK=false
|
||||
if [[ "$WRONG_OWNER" -gt 0 ]]; then
|
||||
find "$SHARE" \( ! -user "$PERMISSIONS_USER" -o ! -group "$PERMISSIONS_GROUP" \) \
|
||||
-exec chown "$PERMISSIONS_OWNER" {} + 2>/dev/null || CHOWN_OK=false
|
||||
fi
|
||||
|
||||
# Apply directory permissions — separate pass for correctness
|
||||
# Directories need execute bit — different from files
|
||||
find "$SHARE" -type d -exec chmod "${PERMISSIONS_DIR_MODE:-755}" {} + \
|
||||
find "$SHARE" -type d ! -perm "${PERMISSIONS_DIR_MODE:-755}" \
|
||||
-exec chmod "${PERMISSIONS_DIR_MODE:-755}" {} + \
|
||||
2>/dev/null || CHMOD_DIR_OK=false
|
||||
|
||||
# Apply file permissions — no execute bit on media files
|
||||
find "$SHARE" -type f -exec chmod "${PERMISSIONS_FILE_MODE:-664}" {} + \
|
||||
2>/dev/null || CHMOD_FILE_OK=false
|
||||
# Ignore "No such file" errors: race condition with volatile dirs (e.g. Emby transcodes)
|
||||
_chmod_errs=$(find "$SHARE" -type f ! -perm "${PERMISSIONS_FILE_MODE:-664}" \
|
||||
-exec chmod "${PERMISSIONS_FILE_MODE:-664}" {} + 2>&1 | \
|
||||
grep -v "No such file or directory" | grep -c "chmod:" || true)
|
||||
[[ "$_chmod_errs" -gt 0 ]] && CHMOD_FILE_OK=false
|
||||
|
||||
if [[ "$CHMOD_DIR_OK" == true && \
|
||||
"$CHMOD_FILE_OK" == true && \
|
||||
@@ -222,7 +331,7 @@ elif [[ ${#FAILED[@]} -gt 0 ]]; then
|
||||
notify "Media permissions failed on $(hostname) — ${FAILED[*]}" \
|
||||
"Media Permissions" "warning"
|
||||
else
|
||||
log "$ICON_DONE Status: done — ${#UPDATED[@]} shares updated"
|
||||
echo "$ICON_DONE Status: done — ${#UPDATED[@]} shares updated"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
|
||||
Executable
+911
@@ -0,0 +1,911 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Play State Sync ============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Syncs watched/played state and resume positions across all configured Emby
|
||||
# and Jellyfin servers. Newest timestamp wins — no data is ever lost.
|
||||
#
|
||||
# Users are matched by name (case-insensitive). If a user exists on some servers
|
||||
# but not others, those servers are skipped for that user — no errors, no partial
|
||||
# syncs from unrelated accounts.
|
||||
#
|
||||
# Items are matched by external provider IDs:
|
||||
# Movies → IMDb ID, then TMDB ID
|
||||
# Episodes → TVDB ID + season + episode number
|
||||
# Audio → MusicBrainz Track ID
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# For each matched item across ≥2 servers:
|
||||
# 1. Compare LastPlayedDate across all servers that have a play record.
|
||||
# 2. The server with the newest LastPlayedDate is authoritative.
|
||||
# 3. Push that server's state (Played, PlayCount, LastPlayedDate,
|
||||
# PlaybackPositionTicks) to every other server.
|
||||
# 4. Servers with no record for that item also receive the state.
|
||||
#
|
||||
# Resume positions (partial plays, not marked Played):
|
||||
# Synced by comparing PlaybackPositionTicks when LastPlayedDate is absent.
|
||||
# The higher tick count wins.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Newest Timestamp Wins
|
||||
# No merge logic, no conflict resolution — the server with the most recent
|
||||
# LastPlayedDate is simply authoritative. Simple rules produce predictable
|
||||
# outcomes users can reason about.
|
||||
#
|
||||
# No Data Loss
|
||||
# The sync only pushes state forward — it never clears a Played flag or
|
||||
# resets a resume position to zero. A watch record on any server always
|
||||
# propagates outward, never disappears.
|
||||
#
|
||||
# Provider ID Matching
|
||||
# Items are matched by external IDs (IMDb, TVDB, MusicBrainz), not by
|
||||
# title or file path. This makes matching robust across library reorganisation,
|
||||
# renames, and multi-server path differences.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# The probe fingerprint is written under STATE_DIR, which is not user-writable.
|
||||
# Without root the fingerprint silently fails to persist and the change probe
|
||||
# never suppresses anything.
|
||||
#
|
||||
# jq Dependency Check
|
||||
# Exits if jq is missing. All API response parsing and the epoch comparisons
|
||||
# depend on it — without jq every comparison would silently evaluate empty.
|
||||
#
|
||||
# PLAY_SYNC_ENABLED Gate
|
||||
# Exits cleanly when disabled; no partial runs.
|
||||
#
|
||||
# PLAY_SYNC_REMOTE Gate
|
||||
# When false, only this host's own servers are synced. Remote hosts are skipped
|
||||
# before any network call is attempted.
|
||||
#
|
||||
# Partnership Gate
|
||||
# Remote hosts are skipped when PARTNERSHIP_ENABLED=false. Local Emby↔Jellyfin
|
||||
# sync still runs — a dormant partnership does not disable local work.
|
||||
#
|
||||
# Tailscale Resolution Guard
|
||||
# A remote host whose Tailscale IP cannot be resolved is skipped rather than
|
||||
# contacted at its literal localhost URL, which would otherwise point the sync
|
||||
# at this host's own server and cross-contaminate state.
|
||||
#
|
||||
# Placeholder Credential Guard
|
||||
# Servers whose API key is empty or still a placeholder are dropped from the
|
||||
# list before any request is made.
|
||||
#
|
||||
# Per-Server Reachability
|
||||
# Unreachable servers are skipped individually; one offline server does not
|
||||
# abort the entire sync.
|
||||
#
|
||||
# User Match Required
|
||||
# A user missing from a server is skipped for that server. State is never
|
||||
# written to an unrelated account that happens to exist there.
|
||||
#
|
||||
# Forward-Only Writes
|
||||
# The sync only pushes state forward — it never clears a Played flag or resets
|
||||
# a resume position. The worst outcome of a bad comparison is a no-op, not
|
||||
# erased watch history.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents concurrent runs racing on the same items during the
|
||||
# 30 minute critical window. --wait switches from skip to wait for manual runs.
|
||||
#
|
||||
# Probe Staleness Ceiling
|
||||
# PLAY_SYNC_PROBE_MAX_AGE_HOURS forces a full comparison regardless of the
|
||||
# hash. Fetches happen every run either way, so the probe can only skip
|
||||
# per-item processing — it can never cause a change to be missed outright.
|
||||
#
|
||||
# Dry Run Support
|
||||
# --dry-run performs all comparisons and writes no state.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_TRANSCODE_SERVERS "Name|URL|APIKey|type" entries per host (emby/jellyfin)
|
||||
# Read directly for every HOST[0-9]+ defined — this script
|
||||
# deliberately does NOT call detect_hosts(), because it needs
|
||||
# every host's servers, not just this one's. Self is identified
|
||||
# by comparing HOST* values against hostname -s.
|
||||
# Remote host URLs have localhost rewritten to their Tailscale IP.
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# PLAY_SYNC_ENABLED Master toggle (default: true)
|
||||
# PLAY_SYNC_REMOTE Sync across all hosts via Tailscale (default: true)
|
||||
# false = local servers only (this host's Emby + Jellyfin)
|
||||
# PLAY_SYNC_TYPES Comma-separated item types to sync (default: Movie,Episode —
|
||||
# Audio excluded, music library too large; favorites handled separately)
|
||||
# PLAY_SYNC_PROBE Skip all per-item processing when no play/resume/favorite
|
||||
# state changed since the last successful run (default: true).
|
||||
# The raw API responses are hashed and compared against the
|
||||
# fingerprint stored in STATE_DIR — fetches still happen every
|
||||
# run, so nothing can be missed.
|
||||
# PLAY_SYNC_PROBE_MAX_AGE_HOURS Force a full comparison when the stored fingerprint is
|
||||
# older than this many hours regardless of the hash (default: 24)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# play_state_sync.sh
|
||||
# Sync all matched users across all configured servers.
|
||||
#
|
||||
# play_state_sync.sh --dry-run
|
||||
# Show what would be synced without writing any state.
|
||||
#
|
||||
# play_state_sync.sh --status
|
||||
# Show configured servers, reachability, and user counts.
|
||||
#
|
||||
# play_state_sync.sh --full
|
||||
# Bypass the change probe — always run the full comparison.
|
||||
#
|
||||
# play_state_sync.sh --wait
|
||||
# Wait for an in-progress run to finish instead of exiting. For manual runs
|
||||
# that would otherwise be skipped by the every-30-minute scheduled pass.
|
||||
#
|
||||
# play_state_sync.sh --log
|
||||
# Verbose output — show each item comparison.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# ── Handle --full and --wait before parse_args ────────────────────────────────
|
||||
FULL_SYNC=false
|
||||
LOCK_MODE="strict"
|
||||
_FILTERED=()
|
||||
for _a in "$@"; do
|
||||
if [[ "$_a" == "--full" ]]; then
|
||||
FULL_SYNC=true
|
||||
elif [[ "$_a" == "--wait" ]]; then
|
||||
LOCK_MODE="wait"
|
||||
else
|
||||
_FILTERED+=("$_a")
|
||||
fi
|
||||
done
|
||||
parse_args "${_FILTERED[@]}"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
# The probe fingerprint lives under STATE_DIR — without root it silently fails to persist
|
||||
# and the change probe can never suppress a run.
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "${PLAY_SYNC_ENABLED:-true}" != "true" ]] && echo "Play state sync disabled" && exit 0
|
||||
|
||||
SYNC_TYPES="${PLAY_SYNC_TYPES:-Movie,Episode}"
|
||||
FAV_TYPES="${PLAY_SYNC_FAV_TYPES:-MusicArtist,MusicAlbum,Movie,Series}"
|
||||
log "$ICON_GEAR Config: types=${SYNC_TYPES} favs=${FAV_TYPES} remote=${PLAY_SYNC_REMOTE:-true} probe=${PLAY_SYNC_PROBE:-true}"
|
||||
|
||||
command -v jq >/dev/null 2>&1 || { error "jq is required but not installed"; exit 1; }
|
||||
|
||||
acquire_lock "$LOCK_MODE"
|
||||
|
||||
# ── Build server list across ALL hosts ────────────────────────────────────────
|
||||
# All HOST*_TRANSCODE_SERVERS arrays are loaded into env by load_config.sh.
|
||||
# For remote hosts, localhost in the URL is rewritten to their Tailscale IP.
|
||||
declare -a SRV_NAME SRV_URL SRV_KEY SRV_TYPE
|
||||
_srv_count=0
|
||||
_my_hostname=$(hostname -s)
|
||||
|
||||
_add_server() {
|
||||
local name="$1" url="$2" key="$3" type="$4"
|
||||
[[ -z "$url" || -z "$key" ]] && return
|
||||
[[ "$key" == "YOUR_API_KEY"* || "$key" == "placeholder"* ]] && return
|
||||
SRV_NAME[$_srv_count]="$name"
|
||||
SRV_URL[$_srv_count]="$url"
|
||||
SRV_KEY[$_srv_count]="$key"
|
||||
SRV_TYPE[$_srv_count]="$type"
|
||||
(( _srv_count++ ))
|
||||
}
|
||||
|
||||
for _varname in $(compgen -v | grep -E '^HOST[0-9]+$' | sort); do
|
||||
_num="${_varname//[^0-9]/}"
|
||||
_host_hostname="${!_varname}"
|
||||
[[ -z "$_host_hostname" ]] && continue
|
||||
|
||||
_is_me=false
|
||||
[[ "${_host_hostname,,}" == "${_my_hostname,,}" ]] && _is_me=true
|
||||
|
||||
# Skip remote hosts when PLAY_SYNC_REMOTE=false
|
||||
if [[ "$_is_me" == false && "${PLAY_SYNC_REMOTE:-true}" != "true" ]]; then
|
||||
log "$_host_hostname — remote sync disabled (PLAY_SYNC_REMOTE=false), skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
# Skip remote hosts when partnership is inactive
|
||||
if [[ "$_is_me" == false && "${PARTNERSHIP_ENABLED:-false}" != "true" ]]; then
|
||||
log "$_host_hostname — partnership inactive (PARTNERSHIP_ENABLED=false), skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
# Resolve Tailscale IP for remote hosts
|
||||
_ts_ip=""
|
||||
if [[ "$_is_me" == false ]]; then
|
||||
_ts_ip=$(resolve_tailscale_ip "$_host_hostname")
|
||||
if [[ -z "$_ts_ip" ]]; then
|
||||
log "$_host_hostname — Tailscale IP not found, skipping"
|
||||
continue
|
||||
fi
|
||||
fi
|
||||
|
||||
# Load this host's TRANSCODE_SERVERS array
|
||||
_srv_arr_name="HOST${_num}_TRANSCODE_SERVERS"
|
||||
eval "_host_entries=(\"\${${_srv_arr_name}[@]}\")"
|
||||
[[ "${#_host_entries[@]}" -eq 0 ]] && continue
|
||||
|
||||
for _entry in "${_host_entries[@]}"; do
|
||||
IFS='|' read -r _name _url _key _type <<< "$_entry"
|
||||
[[ "$_type" == "emby" || "$_type" == "jellyfin" ]] || continue
|
||||
# For remote hosts rewrite localhost/127.0.0.1 → Tailscale IP
|
||||
if [[ "$_is_me" == false ]]; then
|
||||
_url="${_url//localhost/$_ts_ip}"
|
||||
_url="${_url//127.0.0.1/$_ts_ip}"
|
||||
fi
|
||||
_add_server "${_name} (${_host_hostname})" "$_url" "$_key" "$_type"
|
||||
done
|
||||
done
|
||||
|
||||
if [[ "$_srv_count" -lt 2 ]]; then
|
||||
error "Need at least 2 media servers configured — found $_srv_count"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ API Helpers ━━━
|
||||
# ==============================================================================================
|
||||
_api_get() {
|
||||
local url="$1" key="$2" endpoint="$3"
|
||||
curl -sf --max-time 30 \
|
||||
-H "X-Emby-Token: $key" \
|
||||
"${url%/}/${endpoint}" 2>/dev/null
|
||||
}
|
||||
|
||||
_api_post() {
|
||||
local url="$1" key="$2" endpoint="$3" data="${4:-}"
|
||||
if [[ -n "$data" ]]; then
|
||||
curl -sf --max-time 30 -s -o /dev/null -w "%{http_code}" -X POST \
|
||||
-H "X-Emby-Token: $key" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$data" \
|
||||
"${url%/}/${endpoint}" 2>/dev/null
|
||||
else
|
||||
curl -sf --max-time 30 -s -o /dev/null -w "%{http_code}" -X POST \
|
||||
-H "X-Emby-Token: $key" \
|
||||
"${url%/}/${endpoint}" 2>/dev/null
|
||||
fi
|
||||
}
|
||||
|
||||
# Ticks → seconds (1 tick = 100ns, 10_000_000 ticks = 1s)
|
||||
_ticks_to_sec() {
|
||||
echo $(( ${1:-0} / 10000000 ))
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY PLAY STATE SYNC STATUS ━━━━━"
|
||||
echo "$ICON_GEAR Remote sync: ${PLAY_SYNC_REMOTE:-true}"
|
||||
echo "$ICON_GEAR Change probe: ${PLAY_SYNC_PROBE:-true}"
|
||||
echo "$ICON_GEAR Item types: $SYNC_TYPES"
|
||||
echo ""
|
||||
for i in $(seq 0 $(( _srv_count - 1 ))); do
|
||||
echo "$ICON_HOST [${SRV_TYPE[$i]}] ${SRV_NAME[$i]} (${SRV_URL[$i]})"
|
||||
_users=$(_api_get "${SRV_URL[$i]}" "${SRV_KEY[$i]}" "Users" 2>/dev/null | jq -r '.[].Name' 2>/dev/null | wc -l)
|
||||
if [[ "$_users" -gt 0 ]]; then
|
||||
echo " $ICON_DONE Reachable — $_users user(s)"
|
||||
_api_get "${SRV_URL[$i]}" "${SRV_KEY[$i]}" "Users" 2>/dev/null \
|
||||
| jq -r '.[].Name' 2>/dev/null | while read -r n; do echo " · $n"; done
|
||||
else
|
||||
echo " $ICON_ERROR Unreachable or no users"
|
||||
fi
|
||||
done
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Main Sync ━━━
|
||||
# ==============================================================================================
|
||||
echo "━━━ $ICON_SYNC Play State Sync — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no state will be written"
|
||||
[[ "$FULL_SYNC" == true ]] && log "Full sync mode — change probe bypassed"
|
||||
|
||||
START=$(date +%s)
|
||||
TOTAL_SYNCED=0
|
||||
TOTAL_SKIPPED=0
|
||||
TOTAL_ERRORS=0
|
||||
|
||||
# ── Step 1: Fetch users from each server ─────────────────────────────────────
|
||||
declare -A SRV_USERS # idx → JSON array string of users
|
||||
|
||||
log "Fetching users..."
|
||||
for i in $(seq 0 $(( _srv_count - 1 ))); do
|
||||
_resp=$(_api_get "${SRV_URL[$i]}" "${SRV_KEY[$i]}" "Users")
|
||||
if [[ -z "$_resp" ]]; then
|
||||
warn "${SRV_NAME[$i]} — unreachable, skipping"
|
||||
SRV_USERS[$i]=""
|
||||
continue
|
||||
fi
|
||||
SRV_USERS[$i]="$_resp"
|
||||
_count=$(echo "$_resp" | jq 'length' 2>/dev/null || echo 0)
|
||||
log "${SRV_NAME[$i]} — $_count user(s)"
|
||||
done
|
||||
|
||||
# ── Step 2: Build cross-server user map ──────────────────────────────────────
|
||||
# lowercase_name → "server_idx:user_id server_idx:user_id ..."
|
||||
declare -A USER_MAP
|
||||
|
||||
for i in $(seq 0 $(( _srv_count - 1 ))); do
|
||||
[[ -z "${SRV_USERS[$i]}" ]] && continue
|
||||
while IFS=$'\t' read -r uid uname; do
|
||||
[[ -z "$uid" || -z "$uname" ]] && continue
|
||||
lname="${uname,,}"
|
||||
if [[ -n "${USER_MAP[$lname]}" ]]; then
|
||||
USER_MAP[$lname]+=" ${i}:${uid}"
|
||||
else
|
||||
USER_MAP[$lname]="${i}:${uid}"
|
||||
fi
|
||||
done < <(echo "${SRV_USERS[$i]}" | jq -r '.[] | [.Id, .Name] | @tsv' 2>/dev/null)
|
||||
done
|
||||
|
||||
# ── Step 2.3: Fetch play/resume/favorite state for every matched user ────────
|
||||
# Raw responses are cached for the sync passes below and hashed for the change
|
||||
# probe. Fetching costs seconds — the per-item comparison is what costs minutes,
|
||||
# so it only runs when a response actually changed since the last successful run.
|
||||
declare -A RESP_STATE # "lname|si" → deduplicated played+resumable items JSON
|
||||
declare -A RESP_FAV # "lname|si|ftype" → favorites response JSON
|
||||
|
||||
IFS=',' read -ra _fav_type_list <<< "$FAV_TYPES"
|
||||
|
||||
for lname in "${!USER_MAP[@]}"; do
|
||||
read -ra _pairs <<< "${USER_MAP[$lname]}"
|
||||
[[ "${#_pairs[@]}" -lt 2 ]] && continue
|
||||
|
||||
for _pair in "${_pairs[@]}"; do
|
||||
IFS=':' read -r _si _uid <<< "$_pair"
|
||||
|
||||
# All played items — no limit, covers both date-stamped and batch-marked (null date) entries
|
||||
_endpoint="Users/${_uid}/Items?Recursive=true&Fields=ProviderIds,UserData,Type,ParentIndexNumber,IndexNumber,SeriesName&IncludeItemTypes=${SYNC_TYPES}&Filters=IsPlayed&SortBy=DatePlayed&SortOrder=Descending"
|
||||
_resp=$(_api_get "${SRV_URL[$_si]}" "${SRV_KEY[$_si]}" "$_endpoint")
|
||||
if [[ -z "$_resp" ]]; then
|
||||
warn " ${SRV_NAME[$_si]} — failed to fetch items for $lname"
|
||||
RESP_STATE["$lname|$_si"]=""
|
||||
else
|
||||
# Resume positions (not yet marked played)
|
||||
_endpoint2="Users/${_uid}/Items?Recursive=true&Fields=ProviderIds,UserData,Type,ParentIndexNumber,IndexNumber,SeriesName&IncludeItemTypes=${SYNC_TYPES}&SortBy=DatePlayed&SortOrder=Descending&Filters=IsResumable"
|
||||
_resp2=$(_api_get "${SRV_URL[$_si]}" "${SRV_KEY[$_si]}" "$_endpoint2")
|
||||
|
||||
# Combine and deduplicate by Id
|
||||
[[ -z "$_resp2" ]] && _resp2='{"Items":[]}'
|
||||
RESP_STATE["$lname|$_si"]=$(printf '%s\n%s' "$_resp" "$_resp2" | jq -s \
|
||||
'[.[0].Items // [], .[1].Items // []] | add // [] | unique_by(.Id)' 2>/dev/null)
|
||||
fi
|
||||
|
||||
for _ftype in "${_fav_type_list[@]}"; do
|
||||
RESP_FAV["$lname|$_si|$_ftype"]=$(_api_get "${SRV_URL[$_si]}" "${SRV_KEY[$_si]}" \
|
||||
"Users/${_uid}/Items?Recursive=true&IncludeItemTypes=${_ftype}&Filters=IsFavorite&Fields=ProviderIds")
|
||||
done
|
||||
done
|
||||
done
|
||||
|
||||
# ── Step 2.4: Change probe — skip the comparison when nothing changed ────────
|
||||
# Failed fetches hash as empty strings, so reachability transitions also read
|
||||
# as changes and trigger a full pass once the server comes back.
|
||||
PROBE_FILE="$STATE_DIR/play_state_sync_probe"
|
||||
CUR_HASH=$(
|
||||
{
|
||||
while IFS= read -r _k; do
|
||||
printf '%s:' "$_k"; printf '%s' "${RESP_STATE[$_k]}" | md5sum
|
||||
done < <(printf '%s\n' "${!RESP_STATE[@]}" | sort)
|
||||
while IFS= read -r _k; do
|
||||
printf '%s:' "$_k"; printf '%s' "${RESP_FAV[$_k]}" | md5sum
|
||||
done < <(printf '%s\n' "${!RESP_FAV[@]}" | sort)
|
||||
} | md5sum | awk '{print $1}'
|
||||
)
|
||||
|
||||
if [[ "${PLAY_SYNC_PROBE:-true}" == "true" && "$FULL_SYNC" == false && "$DRY_RUN" == false && -f "$PROBE_FILE" ]]; then
|
||||
_prev_hash=$(sed -n '1p' "$PROBE_FILE" 2>/dev/null)
|
||||
_prev_epoch=$(sed -n '2p' "$PROBE_FILE" 2>/dev/null)
|
||||
[[ "$_prev_epoch" =~ ^[0-9]+$ ]] || _prev_epoch=0
|
||||
_probe_max_age=$(( ${PLAY_SYNC_PROBE_MAX_AGE_HOURS:-24} * 3600 ))
|
||||
if [[ "$CUR_HASH" == "$_prev_hash" ]] && (( $(date +%s) - ${_prev_epoch:-0} < _probe_max_age )); then
|
||||
END=$(date +%s)
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY PLAY STATE SYNC SUMMARY ━━━━━"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo "$ICON_DONE No play/resume/favorite changes since last sync — comparison skipped"
|
||||
success "Done ✅"
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Step 2.5: Pre-build provider ID → item ID lookup map ─────────────────────
|
||||
# AnyProviderIdEquals is broken in Jellyfin 10.11+ (ignores the filter entirely).
|
||||
# Pre-fetching all items' provider IDs once and using a local map avoids the broken
|
||||
# per-item API call and is faster overall.
|
||||
declare -A PROV_LOOKUP # "si|tvdb.{id}" | "si|imdb.{id}" | "si|mb.{id}" → item_id
|
||||
|
||||
# Jellyfin's mixed-type query (Movie,Episode,Audio) changes sort order unpredictably,
|
||||
# pushing items to positions far beyond the first pages. Query each type separately so
|
||||
# each list sorts within its own type and items appear where expected.
|
||||
_PROV_PAGE=2000
|
||||
# Combine play state types + fav types for the map build (deduplicated)
|
||||
declare -A _prov_seen_dedup
|
||||
_prov_types=()
|
||||
IFS=',' read -ra _prov_combined <<< "${SYNC_TYPES},${FAV_TYPES}"
|
||||
for _t in "${_prov_combined[@]}"; do
|
||||
[[ -n "${_prov_seen_dedup[$_t]:-}" ]] && continue
|
||||
_prov_seen_dedup[$_t]=1; _prov_types+=("$_t")
|
||||
done
|
||||
unset _prov_seen_dedup _prov_combined
|
||||
|
||||
for _psi in $(seq 0 $(( _srv_count - 1 ))); do
|
||||
[[ -z "${SRV_URL[$_psi]}" ]] && continue
|
||||
log "Building provider ID map for ${SRV_NAME[$_psi]}..."
|
||||
for _prov_type in "${_prov_types[@]}"; do
|
||||
_prov_start=0
|
||||
_prov_total=-1
|
||||
while true; do
|
||||
_page=$(curl -sf --max-time 60 \
|
||||
-H "X-Emby-Token: ${SRV_KEY[$_psi]}" \
|
||||
"${SRV_URL[$_psi]%/}/Items?Recursive=true&IncludeItemTypes=${_prov_type}&Fields=ProviderIds&Limit=${_PROV_PAGE}&StartIndex=${_prov_start}" 2>/dev/null)
|
||||
[[ -z "$_page" ]] && break
|
||||
[[ "$_prov_total" -lt 0 ]] && _prov_total=$(echo "$_page" | jq '.TotalRecordCount // 0' 2>/dev/null || echo 0)
|
||||
_prov_count=$(echo "$_page" | jq '.Items | length' 2>/dev/null || echo 0)
|
||||
[[ "$_prov_count" -eq 0 ]] && break
|
||||
while IFS=$'\t' read -r _pid _ptype _ptvdb _pimdb _ptmdb _pmbtrack _pmbartist _pmbalbum; do
|
||||
[[ "$_pimdb" != "null" && -n "$_pimdb" ]] && PROV_LOOKUP["${_psi}|imdb.${_pimdb}"]="$_pid"
|
||||
[[ "$_ptmdb" != "null" && -n "$_ptmdb" ]] && PROV_LOOKUP["${_psi}|tmdb.${_ptmdb}"]="$_pid"
|
||||
[[ "$_pmbtrack" != "null" && -n "$_pmbtrack" ]] && PROV_LOOKUP["${_psi}|mb.${_pmbtrack}"]="$_pid"
|
||||
[[ "$_pmbartist" != "null" && -n "$_pmbartist" ]] && PROV_LOOKUP["${_psi}|mb.artist.${_pmbartist}"]="$_pid"
|
||||
[[ "$_pmbalbum" != "null" && -n "$_pmbalbum" ]] && PROV_LOOKUP["${_psi}|mb.album.${_pmbalbum}"]="$_pid"
|
||||
# tvdb: namespace by type so Series and Episode IDs don't collide
|
||||
if [[ "$_ptvdb" != "null" && -n "$_ptvdb" ]]; then
|
||||
case "$_ptype" in
|
||||
Series) PROV_LOOKUP["${_psi}|tvdb.series.${_ptvdb}"]="$_pid" ;;
|
||||
Episode) PROV_LOOKUP["${_psi}|tvdb.${_ptvdb}"]="$_pid" ;;
|
||||
*) PROV_LOOKUP["${_psi}|tvdb.${_ptvdb}"]="$_pid" ;;
|
||||
esac
|
||||
fi
|
||||
done < <(echo "$_page" | jq -r '.Items[] | [
|
||||
.Id,
|
||||
.Type,
|
||||
(.ProviderIds.Tvdb // "null"),
|
||||
(.ProviderIds.Imdb // "null"),
|
||||
(.ProviderIds.Tmdb // "null"),
|
||||
(.ProviderIds.MusicBrainzTrackId // "null"),
|
||||
(.ProviderIds.MusicBrainzArtistId // "null"),
|
||||
(.ProviderIds.MusicBrainzAlbumId // "null")
|
||||
] | @tsv' 2>/dev/null)
|
||||
_prov_start=$(( _prov_start + _prov_count ))
|
||||
[[ "$_prov_total" -gt 0 && "$_prov_start" -ge "$_prov_total" ]] && break
|
||||
done
|
||||
done
|
||||
done
|
||||
|
||||
# ── Step 3: Sync per matched user ────────────────────────────────────────────
|
||||
for lname in "${!USER_MAP[@]}"; do
|
||||
read -ra _pairs <<< "${USER_MAP[$lname]}"
|
||||
|
||||
# Skip users only on one server
|
||||
[[ "${#_pairs[@]}" -lt 2 ]] && log " $lname — only on 1 server, skipping" && continue
|
||||
|
||||
echo ""
|
||||
echo "── User: $lname (${#_pairs[@]} server(s)) ──"
|
||||
|
||||
# Build per-server user context
|
||||
declare -A U_IDX U_UID
|
||||
for _pair in "${_pairs[@]}"; do
|
||||
IFS=':' read -r _si _ui <<< "$_pair"
|
||||
U_IDX["$_si"]="$_si"
|
||||
U_UID["$_si"]="$_ui"
|
||||
done
|
||||
|
||||
# ── Fetch played items from each server for this user ────────────────────
|
||||
# Key: provider_id_string → sorted list of (epoch, srv_idx, item_id, play_count, ticks, played)
|
||||
declare -A ITEM_MAP # provider_key → JSON per-server data
|
||||
|
||||
for _si in "${!U_IDX[@]}"; do
|
||||
_combined="${RESP_STATE[$lname|$_si]:-}"
|
||||
[[ -z "$_combined" ]] && continue
|
||||
|
||||
_count=$(echo "$_combined" | jq 'length' 2>/dev/null || echo 0)
|
||||
log " ${SRV_NAME[$_si]} — $_count item(s) with state for $lname"
|
||||
|
||||
# Build item lookup by provider key
|
||||
while IFS=$'\t' read -r iid itype season ep imdb tmdb tvdb mbtrack played ticks lplayed pcount epoch; do
|
||||
# Build canonical provider key
|
||||
_pkey=""
|
||||
case "$itype" in
|
||||
Movie)
|
||||
[[ "$imdb" != "null" && -n "$imdb" ]] && _pkey="imdb:${imdb}"
|
||||
[[ -z "$_pkey" && "$tmdb" != "null" && -n "$tmdb" ]] && _pkey="tmdb:movie:${tmdb}"
|
||||
;;
|
||||
Episode)
|
||||
[[ "$tvdb" != "null" && -n "$tvdb" && "$season" != "null" && "$ep" != "null" ]] && \
|
||||
_pkey="tvdb:ep:${tvdb}:s${season}e${ep}"
|
||||
;;
|
||||
Audio)
|
||||
[[ "$mbtrack" != "null" && -n "$mbtrack" ]] && _pkey="mb:track:${mbtrack}"
|
||||
;;
|
||||
esac
|
||||
[[ -z "$_pkey" ]] && continue
|
||||
|
||||
_epoch="${epoch:-0}"
|
||||
_entry="${_si}|${iid}|${played}|${pcount}|${ticks}|${_epoch}|${lplayed}"
|
||||
|
||||
if [[ -n "${ITEM_MAP[$_pkey]}" ]]; then
|
||||
ITEM_MAP[$_pkey]+=$'\n'"$_entry"
|
||||
else
|
||||
ITEM_MAP[$_pkey]="$_entry"
|
||||
fi
|
||||
|
||||
done < <(echo "$_combined" | jq -r '.[] | [
|
||||
.Id,
|
||||
.Type,
|
||||
(.ParentIndexNumber // "null" | tostring),
|
||||
(.IndexNumber // "null" | tostring),
|
||||
(.ProviderIds.Imdb // "null"),
|
||||
(.ProviderIds.Tmdb // "null"),
|
||||
(.ProviderIds.Tvdb // "null"),
|
||||
(.ProviderIds.MusicBrainzTrackId // "null"),
|
||||
(.UserData.Played // false | tostring),
|
||||
(.UserData.PlaybackPositionTicks // 0 | tostring),
|
||||
(.UserData.LastPlayedDate // "null"),
|
||||
(.UserData.PlayCount // 0 | tostring),
|
||||
((.UserData.LastPlayedDate // null) | if . == null then "0"
|
||||
else (((sub("\\.[0-9]+"; "") | sub("[+-][0-9]{2}:?[0-9]{2}$"; "Z")
|
||||
| if endswith("Z") then . else . + "Z" end
|
||||
| fromdateiso8601)? // 0) | tostring) end)
|
||||
] | @tsv' 2>/dev/null)
|
||||
done
|
||||
|
||||
# ── Compare and sync ─────────────────────────────────────────────────────
|
||||
for _pkey in "${!ITEM_MAP[@]}"; do
|
||||
# Collect all server entries for this item
|
||||
declare -A E_EPOCH E_PLAYED E_PCOUNT E_TICKS E_IID E_LPLAYED E_SIDX
|
||||
_has_entries=false
|
||||
|
||||
while IFS='|' read -r _si _iid _played _pcount _ticks _epoch _lplayed; do
|
||||
[[ -z "$_si" ]] && continue
|
||||
E_SIDX[$_si]="$_si"
|
||||
E_IID[$_si]="$_iid"
|
||||
E_PLAYED[$_si]="$_played"
|
||||
E_PCOUNT[$_si]="$_pcount"
|
||||
E_TICKS[$_si]="$_ticks"
|
||||
E_EPOCH[$_si]="$_epoch"
|
||||
E_LPLAYED[$_si]="$_lplayed"
|
||||
_has_entries=true
|
||||
done <<< "${ITEM_MAP[$_pkey]}"
|
||||
|
||||
[[ "$_has_entries" == false ]] && continue
|
||||
|
||||
# Find the authoritative server: newest LastPlayedDate epoch
|
||||
# Tie-break: higher PlayCount, then higher Ticks, then Played=true
|
||||
# Init at -1 so servers with epoch=0 (batch-marks with null LastPlayedDate) can win
|
||||
_auth_si=""
|
||||
_auth_epoch=-1
|
||||
_auth_pcount=-1
|
||||
_auth_ticks=-1
|
||||
_auth_pf="false"
|
||||
|
||||
for _si in "${!E_SIDX[@]}"; do
|
||||
_e="${E_EPOCH[$_si]:-0}"
|
||||
_pc="${E_PCOUNT[$_si]:-0}"
|
||||
_tk="${E_TICKS[$_si]:-0}"
|
||||
_pf="${E_PLAYED[$_si]:-false}"
|
||||
# Played=true is the primary key — a played server always beats a non-played server
|
||||
# regardless of epoch. A resumable item with a newer LastPlayedDate must not become
|
||||
# authority over a played item, as pushing resume ticks to a played server resets
|
||||
# the played status on some Emby/Jellyfin versions.
|
||||
if ( [[ "$_pf" == "true" ]] && [[ "$_auth_pf" != "true" ]] ) || \
|
||||
( [[ "$_pf" == "$_auth_pf" ]] && [[ "$_e" -gt "$_auth_epoch" ]] ) || \
|
||||
( [[ "$_pf" == "$_auth_pf" ]] && [[ "$_e" -eq "$_auth_epoch" ]] && [[ "$_pc" -gt "$_auth_pcount" ]] ) || \
|
||||
( [[ "$_pf" == "$_auth_pf" ]] && [[ "$_e" -eq "$_auth_epoch" ]] && [[ "$_pc" -eq "$_auth_pcount" ]] && [[ "$_tk" -gt "$_auth_ticks" ]] ); then
|
||||
_auth_si="$_si"
|
||||
_auth_epoch="$_e"
|
||||
_auth_pcount="$_pc"
|
||||
_auth_ticks="$_tk"
|
||||
_auth_pf="$_pf"
|
||||
fi
|
||||
done
|
||||
|
||||
[[ -z "$_auth_si" ]] && continue
|
||||
|
||||
_auth_played="${E_PLAYED[$_auth_si]}"
|
||||
_auth_lplayed="${E_LPLAYED[$_auth_si]}"
|
||||
_auth_pcount="${E_PCOUNT[$_auth_si]}"
|
||||
_auth_ticks="${E_TICKS[$_auth_si]}"
|
||||
|
||||
# Push to servers with older state OR no state at all
|
||||
for _si in "${!U_IDX[@]}"; do
|
||||
[[ "$_si" == "$_auth_si" ]] && continue
|
||||
_their_epoch="${E_EPOCH[$_si]:-0}"
|
||||
_their_played="${E_PLAYED[$_si]:-false}"
|
||||
|
||||
# Skip if they already have the same/newer state
|
||||
_skip=false
|
||||
if [[ "$_auth_played" == "true" ]]; then
|
||||
# Both servers already have this played — nothing to propagate regardless of dates.
|
||||
# Date-based comparison caused a ping-pong: syncing without a DatePlayed param lets
|
||||
# the target server stamp the current time, making it the new authority next cycle.
|
||||
[[ "$_their_played" == "true" ]] && _skip=true
|
||||
else
|
||||
# Resume only: skip if target already has same or more ticks.
|
||||
# Allow 5-second tolerance (50_000_000 ticks) — Emby may round tick values slightly
|
||||
# differently on read, causing an exact-match check to miss and re-sync every cycle.
|
||||
_tick_gap=$(( ${_auth_ticks:-0} - ${E_TICKS[$_si]:-0} ))
|
||||
[[ "$_tick_gap" -le 50000000 && "${E_TICKS[$_si]:-0}" -gt 0 && "$_their_played" == "false" ]] && _skip=true
|
||||
fi
|
||||
if [[ "$_skip" == true ]]; then
|
||||
log " SKIP $_pkey → ${SRV_NAME[$_si]} already up to date"
|
||||
(( TOTAL_SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
_uid="${U_UID[$_si]}"
|
||||
_iid="${E_IID[$_si]:-}" # might not exist on this server yet
|
||||
|
||||
# Find item ID on target server by provider key if not in our map
|
||||
if [[ -z "$_iid" ]]; then
|
||||
_ptype="${_pkey%%:*}"
|
||||
case "$_ptype" in
|
||||
imdb)
|
||||
_pval="${_pkey#imdb:}"
|
||||
_iid="${PROV_LOOKUP[${_si}|imdb.${_pval}]:-}"
|
||||
;;
|
||||
tmdb)
|
||||
_pval="${_pkey#tmdb:movie:}"
|
||||
_iid="${PROV_LOOKUP[${_si}|tmdb.${_pval}]:-}"
|
||||
;;
|
||||
tvdb)
|
||||
# pkey format: tvdb:ep:{tvdb_id}:s{season}e{ep}
|
||||
_tvdb_num="${_pkey#tvdb:ep:}"; _tvdb_num="${_tvdb_num%%:*}"
|
||||
_iid="${PROV_LOOKUP[${_si}|tvdb.${_tvdb_num}]:-}"
|
||||
;;
|
||||
mb)
|
||||
_pval="${_pkey#mb:track:}"
|
||||
_iid="${PROV_LOOKUP[${_si}|mb.${_pval}]:-}"
|
||||
;;
|
||||
esac
|
||||
[[ -z "$_iid" ]] && log " SKIP $_pkey → ${SRV_NAME[$_si]} item not found on server" && continue
|
||||
fi
|
||||
|
||||
log " SYNC $_pkey → ${SRV_NAME[$_si]} (auth: ${SRV_NAME[$_auth_si]}, epoch: $_auth_epoch)"
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
echo " DRY RUN: would sync $_pkey → ${SRV_NAME[$_si]} user=$lname played=$_auth_played date=$_auth_lplayed"
|
||||
(( TOTAL_SYNCED++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# Write state to target server
|
||||
if [[ "$_auth_played" == "true" ]]; then
|
||||
# Mark as played with date
|
||||
_date_param=""
|
||||
if [[ "$_auth_lplayed" != "null" && -n "$_auth_lplayed" ]]; then
|
||||
# PlayedItems expects DatePlayed in yyyyMMddHHmmss (no separators, no timezone)
|
||||
_lp="${_auth_lplayed%.*}"
|
||||
_lp="${_lp%Z}"
|
||||
_date_param="?DatePlayed=${_lp//[-T:]/}"
|
||||
fi
|
||||
_http=$(_api_post "${SRV_URL[$_si]}" "${SRV_KEY[$_si]}" \
|
||||
"Users/${_uid}/PlayedItems/${_iid}${_date_param}")
|
||||
if [[ "$_http" == "200" || "$_http" == "201" ]]; then
|
||||
success " ✓ $_pkey → ${SRV_NAME[$_si]} marked played"
|
||||
(( TOTAL_SYNCED++ ))
|
||||
else
|
||||
warn " ✗ $_pkey → ${SRV_NAME[$_si]} failed (HTTP ${_http:-err})"
|
||||
(( TOTAL_ERRORS++ ))
|
||||
fi
|
||||
else
|
||||
# Sync resume position only — never push ticks to a server that already
|
||||
# has this item marked played. Writing ticks via UserData can reset Played=false.
|
||||
if [[ "${E_PLAYED[$_si]:-false}" == "true" ]]; then
|
||||
log " SKIP $_pkey → ${SRV_NAME[$_si]} already played, not overwriting with resume ticks"
|
||||
(( TOTAL_SKIPPED++ ))
|
||||
continue
|
||||
fi
|
||||
_ticks_int=$(( ${_auth_ticks:-0} ))
|
||||
if [[ "$_ticks_int" -gt 0 ]]; then
|
||||
_payload="{\"PlaybackPositionTicks\":${_ticks_int}}"
|
||||
_http=$(_api_post "${SRV_URL[$_si]}" "${SRV_KEY[$_si]}" \
|
||||
"Users/${_uid}/Items/${_iid}/UserData" "$_payload")
|
||||
if [[ "$_http" == "200" || "$_http" == "204" ]]; then
|
||||
success " ✓ $_pkey → ${SRV_NAME[$_si]} resume synced ($(_ticks_to_sec "$_ticks_int")s)"
|
||||
(( TOTAL_SYNCED++ ))
|
||||
else
|
||||
warn " ✗ $_pkey → ${SRV_NAME[$_si]} resume sync failed (HTTP ${_http:-err})"
|
||||
(( TOTAL_ERRORS++ ))
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
unset E_EPOCH E_PLAYED E_PCOUNT E_TICKS E_IID E_LPLAYED E_SIDX
|
||||
done
|
||||
|
||||
unset ITEM_MAP U_IDX U_UID
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Favorite Sync ━━━
|
||||
# ==============================================================================================
|
||||
# Union semantics — if favorited on any server, sync to all others. Never unmarks.
|
||||
# Covers MusicArtist, MusicAlbum, Movie, Series (configured via PLAY_SYNC_FAV_TYPES).
|
||||
FAV_TOTAL_SYNCED=0
|
||||
FAV_TOTAL_SKIPPED=0
|
||||
FAV_TOTAL_ERRORS=0
|
||||
|
||||
for lname in "${!USER_MAP[@]}"; do
|
||||
read -ra _pairs <<< "${USER_MAP[$lname]}"
|
||||
[[ "${#_pairs[@]}" -lt 2 ]] && continue
|
||||
|
||||
declare -A UF_IDX UF_UID
|
||||
for _pair in "${_pairs[@]}"; do
|
||||
IFS=':' read -r _si _ui <<< "$_pair"
|
||||
UF_IDX["$_si"]="$_si"
|
||||
UF_UID["$_si"]="$_ui"
|
||||
done
|
||||
|
||||
declare -A FAV_MAP # provider_key → space-separated "si:iid" pairs
|
||||
|
||||
for _si in "${!UF_IDX[@]}"; do
|
||||
_uid="${UF_UID[$_si]}"
|
||||
for _ftype in "${_fav_type_list[@]}"; do
|
||||
_resp="${RESP_FAV[$lname|$_si|$_ftype]:-}"
|
||||
[[ -z "$_resp" ]] && continue
|
||||
_fcount=$(echo "$_resp" | jq '.Items | length' 2>/dev/null || echo 0)
|
||||
[[ "$_fcount" -eq 0 ]] && continue
|
||||
log " ${SRV_NAME[$_si]} — $_fcount ${_ftype} favorite(s) for $lname"
|
||||
|
||||
while IFS=$'\t' read -r _iid _itype _pimdb _ptmdb _ptvdb _pmbartist _pmbalbum _pmbtrack; do
|
||||
_pkey=""
|
||||
case "$_itype" in
|
||||
Movie)
|
||||
[[ "$_pimdb" != "null" && -n "$_pimdb" ]] && _pkey="imdb:${_pimdb}"
|
||||
[[ -z "$_pkey" && "$_ptmdb" != "null" && -n "$_ptmdb" ]] && _pkey="tmdb:movie:${_ptmdb}"
|
||||
;;
|
||||
Series)
|
||||
[[ "$_ptvdb" != "null" && -n "$_ptvdb" ]] && _pkey="tvdb:series:${_ptvdb}"
|
||||
;;
|
||||
MusicArtist)
|
||||
[[ "$_pmbartist" != "null" && -n "$_pmbartist" ]] && _pkey="mb:artist:${_pmbartist}"
|
||||
;;
|
||||
MusicAlbum)
|
||||
[[ "$_pmbalbum" != "null" && -n "$_pmbalbum" ]] && _pkey="mb:album:${_pmbalbum}"
|
||||
;;
|
||||
Audio)
|
||||
[[ "$_pmbtrack" != "null" && -n "$_pmbtrack" ]] && _pkey="mb:track:${_pmbtrack}"
|
||||
;;
|
||||
esac
|
||||
[[ -z "$_pkey" ]] && continue
|
||||
|
||||
if [[ -n "${FAV_MAP[$_pkey]:-}" ]]; then
|
||||
FAV_MAP[$_pkey]+=" ${_si}:${_iid}"
|
||||
else
|
||||
FAV_MAP[$_pkey]="${_si}:${_iid}"
|
||||
fi
|
||||
done < <(echo "$_resp" | jq -r '.Items[] | [
|
||||
.Id, .Type,
|
||||
(.ProviderIds.Imdb // "null"),
|
||||
(.ProviderIds.Tmdb // "null"),
|
||||
(.ProviderIds.Tvdb // "null"),
|
||||
(.ProviderIds.MusicBrainzArtistId // "null"),
|
||||
(.ProviderIds.MusicBrainzAlbumId // "null"),
|
||||
(.ProviderIds.MusicBrainzTrackId // "null")
|
||||
] | @tsv' 2>/dev/null)
|
||||
done
|
||||
done
|
||||
|
||||
for _pkey in "${!FAV_MAP[@]}"; do
|
||||
declare -A _fhave _fiid
|
||||
for _entry in ${FAV_MAP[$_pkey]}; do
|
||||
IFS=':' read -r _si _iid <<< "$_entry"
|
||||
_fhave[$_si]=true
|
||||
_fiid[$_si]="$_iid"
|
||||
done
|
||||
|
||||
for _si in "${!UF_IDX[@]}"; do
|
||||
[[ "${_fhave[$_si]:-false}" == "true" ]] && (( FAV_TOTAL_SKIPPED++ )) && continue
|
||||
_uid="${UF_UID[$_si]}"
|
||||
|
||||
# Look up item ID on target server via PROV_LOOKUP
|
||||
_iid=""
|
||||
_ptype="${_pkey%%:*}"
|
||||
case "$_ptype" in
|
||||
imdb) _iid="${PROV_LOOKUP[${_si}|imdb.${_pkey#imdb:}]:-}" ;;
|
||||
tmdb) _iid="${PROV_LOOKUP[${_si}|tmdb.${_pkey#tmdb:movie:}]:-}" ;;
|
||||
tvdb) _iid="${PROV_LOOKUP[${_si}|tvdb.series.${_pkey#tvdb:series:}]:-}" ;;
|
||||
mb)
|
||||
_mbsub="${_pkey#mb:}"
|
||||
case "${_mbsub%%:*}" in
|
||||
artist) _iid="${PROV_LOOKUP[${_si}|mb.artist.${_mbsub#artist:}]:-}" ;;
|
||||
album) _iid="${PROV_LOOKUP[${_si}|mb.album.${_mbsub#album:}]:-}" ;;
|
||||
track) _iid="${PROV_LOOKUP[${_si}|mb.${_mbsub#track:}]:-}" ;;
|
||||
esac
|
||||
;;
|
||||
esac
|
||||
|
||||
if [[ -z "$_iid" ]]; then
|
||||
log " SKIP FAV $_pkey → ${SRV_NAME[$_si]} not in library"
|
||||
continue
|
||||
fi
|
||||
|
||||
log " FAV $_pkey → ${SRV_NAME[$_si]} ($lname)"
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
_http=$(_api_post "${SRV_URL[$_si]}" "${SRV_KEY[$_si]}" \
|
||||
"Users/${_uid}/FavoriteItems/${_iid}")
|
||||
if [[ "$_http" == "200" || "$_http" == "201" ]]; then
|
||||
success " ✓ FAV $_pkey → ${SRV_NAME[$_si]}"
|
||||
(( FAV_TOTAL_SYNCED++ ))
|
||||
else
|
||||
warn " ✗ FAV $_pkey → ${SRV_NAME[$_si]} failed (HTTP ${_http:-err})"
|
||||
(( FAV_TOTAL_ERRORS++ ))
|
||||
fi
|
||||
else
|
||||
echo " DRY RUN: would favorite $_pkey → ${SRV_NAME[$_si]} user=$lname"
|
||||
(( FAV_TOTAL_SYNCED++ ))
|
||||
fi
|
||||
done
|
||||
|
||||
unset _fhave _fiid
|
||||
done
|
||||
|
||||
unset FAV_MAP UF_IDX UF_UID
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
END=$(date +%s)
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY PLAY STATE SYNC SUMMARY ━━━━━"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo "$ICON_DONE Synced: $TOTAL_SYNCED"
|
||||
[[ "$TOTAL_SKIPPED" -gt 0 ]] && echo "$ICON_RUNNING Skipped: $TOTAL_SKIPPED (already current)"
|
||||
[[ "$TOTAL_ERRORS" -gt 0 ]] && echo "$ICON_ERROR Errors: $TOTAL_ERRORS"
|
||||
echo ""
|
||||
echo " Favorites — synced: $FAV_TOTAL_SYNCED skipped: $FAV_TOTAL_SKIPPED errors: $FAV_TOTAL_ERRORS"
|
||||
|
||||
TOTAL_ALL_ERRORS=$(( TOTAL_ERRORS + FAV_TOTAL_ERRORS ))
|
||||
|
||||
# Fingerprint is the PRE-sync state — our own writes above changed the targets,
|
||||
# so the next run does one more full pass and then settles into probe skips.
|
||||
# Re-hashing post-sync instead would swallow plays that landed mid-run.
|
||||
if [[ "$DRY_RUN" == false && "$TOTAL_ALL_ERRORS" -eq 0 ]]; then
|
||||
printf '%s\n%s\n' "$CUR_HASH" "$START" > "$PROBE_FILE" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes written"
|
||||
exit 0
|
||||
elif [[ "$TOTAL_ALL_ERRORS" -eq 0 ]]; then
|
||||
success "Done ✅"
|
||||
exit 0
|
||||
else
|
||||
warn "Done with $TOTAL_ALL_ERRORS error(s)"
|
||||
exit 1
|
||||
fi
|
||||
@@ -1,586 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Radarr Cleanup =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Delete orphaned movie files not tracked by Radarr. Queries the API for all
|
||||
# tracked movie file paths, walks the library on disk, and removes anything
|
||||
# untracked that is old enough to be past the import window. Triggers an Emby
|
||||
# library clean after each deletion run so ghost entries disappear immediately.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Every file encountered on disk is classified into one of five categories:
|
||||
#
|
||||
# TRACKED — Radarr API knows this exact path → leave it alone
|
||||
# PROTECTED — matches RADARR_PROTECTED_PATTERNS → never delete
|
||||
# ORPHAN — video file, not tracked, older than RADARR_ORPHAN_AGE → delete
|
||||
# JUNK — not a video extension, not protected → delete regardless of age
|
||||
# RECENT — not tracked, under RADARR_ORPHAN_AGE → skip (may be mid-import)
|
||||
#
|
||||
# Radarr generates movie artwork (*.jpg), metadata (*.nfo), and manages subtitles
|
||||
# (*.srt, *.sub, *.ass) but does NOT include these in its tracked file API response.
|
||||
# Without PROTECTED classification these would be deleted — breaking Radarr and
|
||||
# Emby metadata display.
|
||||
#
|
||||
# After deletions: notify_emby_scan() triggers Emby "Clean Missing Files" task.
|
||||
# Emby removes ghost entries immediately — no user-facing file-not-found errors.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Six gates — ALL must pass before any file is touched:
|
||||
# 1. Container running and not starting/unhealthy
|
||||
# 2. API reachable
|
||||
# 3. API version matches RADARR_VERSION_MAJOR in master.conf
|
||||
# 4. Movie count > 0
|
||||
# 5. Tracked file count > 0
|
||||
# 6. Deletion size < RADARR_MAX_DELETE_GB — or --i-know-what-im-doing required
|
||||
#
|
||||
# acquire_lock "wait" — large scans take time, wait for previous run to finish
|
||||
# jq + curl validation — exits if either tool missing
|
||||
# DOCKER_TIMEOUT — container checks protected against daemon hangs
|
||||
# notify_emby_scan() — triggers Emby clean after deletion
|
||||
# validate_unraid_cmd — notify script validated before use
|
||||
# Silent by default — orphans/junk warn(), clean library logs silently
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_RADARR_URL / HOST*_RADARR_API_KEY / HOST*_RADARR_MOVIES_ROOT
|
||||
# HOST*_RADARR_PATH_MAP — container path → host path translation
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# RADARR_ORPHAN_AGE — days before untracked file eligible for deletion
|
||||
# RADARR_MAX_DELETE_GB — require --i-know-what-im-doing above this
|
||||
# RADARR_EXTENSIONS — video file extensions for orphan classification
|
||||
# RADARR_PROTECTED_PATTERNS — file patterns never deleted
|
||||
# RADARR_VERSION_MAJOR — expected Radarr major version for API safety check
|
||||
# RADARR_IMPORT_SCAN_TIMEOUT — seconds to wait for pre-flight import scan (default 600)
|
||||
# ARR_CLEANUP_STATS — stats file path (read by coffee report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# radarr_cleanup.sh — normal run
|
||||
# radarr_cleanup.sh --dry-run — preview, no deletions
|
||||
# radarr_cleanup.sh --log — verbose output
|
||||
# radarr_cleanup.sh --status — show config and exit
|
||||
# radarr_cleanup.sh --i-know-what-im-doing — bypass size threshold
|
||||
# radarr_cleanup.sh --i-know-what-im-doing --skip-strike-list — NUCLEAR MODE
|
||||
#
|
||||
# NUCLEAR MODE: both flags bypass age check AND size threshold. User accepts full
|
||||
# responsibility — the flag name is long and annoying by design.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# ── Special flag pre-processing ───────────────────────────────────────────────────────────────
|
||||
I_KNOW=false
|
||||
SKIP_STRIKES=false
|
||||
FILTERED_ARGS=()
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--i-know-what-im-doing) I_KNOW=true ;;
|
||||
--skip-strike-list) SKIP_STRIKES=true ;;
|
||||
*) FILTERED_ARGS+=("$arg") ;;
|
||||
esac
|
||||
done
|
||||
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
|
||||
# ── Nuclear mode warning ──────────────────────────────────────────────────────────────────────
|
||||
if [[ "$I_KNOW" == true ]] && [[ "$SKIP_STRIKES" == true ]] && [[ "$DRY_RUN" != true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo "⚠️ WARNING — NUCLEAR MODE ACTIVE"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo " Flags: --i-know-what-im-doing --skip-strike-list"
|
||||
echo " Strike system: BYPASSED — deletes on first pass"
|
||||
echo " Size threshold: BYPASSED — no GB limit"
|
||||
echo " Data recovery: NOT POSSIBLE after deletion"
|
||||
echo ""
|
||||
echo " Review --dry-run output before proceeding."
|
||||
echo " You have 10 seconds to cancel (Ctrl+C)..."
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
sleep 10
|
||||
echo " Proceeding..."
|
||||
echo ""
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for Radarr API calls"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
notify "Radarr cleanup failed on $(hostname) — jq not installed" "Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock "wait"
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases RADARR_URL, RADARR_API_KEY, RADARR_MOVIES_ROOT
|
||||
detect_hosts
|
||||
|
||||
DOCKER_TIMEOUT=15
|
||||
RADARR_CONTAINER="Radarr"
|
||||
|
||||
# Build path map from MY_ID's Radarr path map
|
||||
declare -A ARR_PATH_MAP
|
||||
local_path_map_var="${MY_ID}_RADARR_PATH_MAP"
|
||||
eval "for key in \"\${!${local_path_map_var}[@]}\"; do
|
||||
ARR_PATH_MAP[\"\$key\"]=\"\${${local_path_map_var}[\$key]}\"
|
||||
done"
|
||||
|
||||
require_var RADARR_URL
|
||||
require_var RADARR_API_KEY
|
||||
require_var RADARR_MOVIES_ROOT
|
||||
|
||||
if [[ ! -d "$RADARR_MOVIES_ROOT" ]]; then
|
||||
error "Movies root not found: $RADARR_MOVIES_ROOT"
|
||||
notify "Radarr cleanup failed on $(hostname) — movies root not found: $RADARR_MOVIES_ROOT" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $RADARR_URL"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be deleted"
|
||||
[[ "$I_KNOW" == true ]] && warn "OVERRIDE — --i-know-what-im-doing active"
|
||||
[[ "$SKIP_STRIKES" == true ]] && warn "OVERRIDE — --skip-strike-list active — age check bypassed"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Radarr URL: $RADARR_URL"
|
||||
echo "$ICON_GEAR Movies root: $RADARR_MOVIES_ROOT"
|
||||
echo "$ICON_TIME Orphan age: ${RADARR_ORPHAN_AGE} days"
|
||||
echo "$ICON_GEAR Max delete: ${RADARR_MAX_DELETE_GB}GB (requires --i-know-what-im-doing)"
|
||||
echo "$ICON_GEAR Radarr ver: v${RADARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Extensions: ${RADARR_EXTENSIONS[*]}"
|
||||
echo "$ICON_GEAR Protected patterns: ${RADARR_PROTECTED_PATTERNS[*]}"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "$ICON_GEAR I know: $I_KNOW"
|
||||
echo "$ICON_GEAR Skip strikes: $SKIP_STRIKES"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 1 — Container Health ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
|
||||
CONTAINER_RUNNING=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Running}}' "$RADARR_CONTAINER" 2>/dev/null)
|
||||
if [[ "$CONTAINER_RUNNING" != "true" ]]; then
|
||||
error "$RADARR_CONTAINER is not running — aborting"
|
||||
notify "Radarr cleanup aborted on $(hostname) — container not running" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
CONTAINER_HEALTH=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Health.Status}}' "$RADARR_CONTAINER" 2>/dev/null)
|
||||
case "$CONTAINER_HEALTH" in
|
||||
healthy) info "$RADARR_CONTAINER is healthy" ;;
|
||||
"") info "$RADARR_CONTAINER has no health check — proceeding" ;;
|
||||
starting)
|
||||
error "$RADARR_CONTAINER is still starting — aborting"
|
||||
notify "Radarr cleanup aborted on $(hostname) — container still starting" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1 ;;
|
||||
unhealthy)
|
||||
error "$RADARR_CONTAINER is unhealthy — aborting"
|
||||
notify "Radarr cleanup aborted on $(hostname) — container unhealthy" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1 ;;
|
||||
*) warn "$RADARR_CONTAINER health: $CONTAINER_HEALTH — proceeding with caution" ;;
|
||||
esac
|
||||
|
||||
info "Safety layer 1 passed — container healthy"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
radarr_api() {
|
||||
local endpoint="$1"
|
||||
local response http_code body
|
||||
|
||||
response=$(curl -sf \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${RADARR_URL}/api/v3/${endpoint}" 2>/dev/null)
|
||||
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
|
||||
if [[ "$http_code" != "200" ]]; then
|
||||
error "Radarr API HTTP $http_code for: $endpoint"
|
||||
return 1
|
||||
fi
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
is_video_file() {
|
||||
local ext="${1##*.}"
|
||||
ext="${ext,,}"
|
||||
for valid_ext in "${RADARR_EXTENSIONS[@]}"; do
|
||||
[[ "$ext" == "$valid_ext" ]] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
is_protected_file() {
|
||||
local filename
|
||||
filename=$(basename "$1")
|
||||
for pattern in "${RADARR_PROTECTED_PATTERNS[@]}"; do
|
||||
# shellcheck disable=SC2254
|
||||
case "$filename" in
|
||||
$pattern) return 0 ;;
|
||||
esac
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
format_bytes() {
|
||||
local bytes=$1
|
||||
if (( bytes > 1073741824 )); then
|
||||
awk "BEGIN {printf \"%.1fGB\", $bytes / 1073741824}"
|
||||
elif (( bytes > 1048576 )); then
|
||||
awk "BEGIN {printf \"%.1fMB\", $bytes / 1048576}"
|
||||
else
|
||||
echo "${bytes}B"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Pre-flight: Radarr Import Scan ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Pre-flight: Radarr Import Scan ━━━"
|
||||
|
||||
# Reverse-lookup container path from path map so Radarr gets its own path, not the host path
|
||||
RADARR_CONTAINER_ROOT=""
|
||||
for _cp in "${!ARR_PATH_MAP[@]}"; do
|
||||
if [[ "${ARR_PATH_MAP[$_cp]}" == "$RADARR_MOVIES_ROOT" ]]; then
|
||||
RADARR_CONTAINER_ROOT="$_cp"
|
||||
break
|
||||
fi
|
||||
done
|
||||
unset _cp
|
||||
|
||||
if [[ -n "$RADARR_CONTAINER_ROOT" ]]; then
|
||||
info "Triggering DownloadedMoviesScan on: $RADARR_CONTAINER_ROOT"
|
||||
SCAN_PAYLOAD="{\"name\": \"DownloadedMoviesScan\", \"path\": \"$RADARR_CONTAINER_ROOT\"}"
|
||||
else
|
||||
info "No path map match — triggering DownloadedMoviesScan (all root folders)"
|
||||
SCAN_PAYLOAD='{"name": "DownloadedMoviesScan"}'
|
||||
fi
|
||||
|
||||
SCAN_RESPONSE=$(curl -sf --max-time 30 -X POST \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$SCAN_PAYLOAD" \
|
||||
"${RADARR_URL}/api/v3/command" 2>/dev/null)
|
||||
|
||||
SCAN_CMD_ID=$(echo "$SCAN_RESPONSE" | jq -r '.id // empty' 2>/dev/null)
|
||||
|
||||
if [[ -z "$SCAN_CMD_ID" ]]; then
|
||||
warn "Could not trigger import scan — proceeding without pre-flight"
|
||||
else
|
||||
info "Import scan queued (command ID: $SCAN_CMD_ID) — waiting for completion..."
|
||||
POLL_TIMEOUT=${RADARR_IMPORT_SCAN_TIMEOUT:-600}
|
||||
POLLED=0
|
||||
while [[ "$POLLED" -lt "$POLL_TIMEOUT" ]]; do
|
||||
SCAN_STATUS=$(curl -sf --max-time 10 \
|
||||
-H "X-Api-Key: $RADARR_API_KEY" \
|
||||
"${RADARR_URL}/api/v3/command/${SCAN_CMD_ID}" 2>/dev/null | \
|
||||
jq -r '.status // empty' 2>/dev/null)
|
||||
case "$SCAN_STATUS" in
|
||||
completed) info "Import scan complete ✅"; break ;;
|
||||
failed) warn "Import scan reported failed — proceeding anyway"; break ;;
|
||||
esac
|
||||
sleep 10
|
||||
(( POLLED += 10 ))
|
||||
[[ $(( POLLED % 60 )) -eq 0 ]] && log " Still scanning... (${POLLED}s elapsed)"
|
||||
done
|
||||
[[ "$POLLED" -ge "$POLL_TIMEOUT" ]] && \
|
||||
warn "Import scan timed out after ${POLL_TIMEOUT}s — proceeding anyway"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Radarr Tracked Files ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching Radarr Tracked Files ━━━"
|
||||
|
||||
# Safety Layer 2 — API reachability
|
||||
if ! check_api "$RADARR_URL" "Radarr" 10; then
|
||||
notify "Radarr cleanup aborted on $(hostname) — API unreachable" "Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Safety Layer 3 — API version check
|
||||
check_arr_version "$RADARR_URL" "$RADARR_API_KEY" "v3" "$RADARR_VERSION_MAJOR" "Radarr" || exit 1
|
||||
|
||||
info "Querying Radarr API..."
|
||||
|
||||
# Fetch all movies
|
||||
MOVIES_RESPONSE=$(radarr_api "movie") || {
|
||||
error "Failed to fetch movies from Radarr"
|
||||
notify "Radarr cleanup failed on $(hostname) — could not fetch movies" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
}
|
||||
|
||||
MOVIE_IDS=$(echo "$MOVIES_RESPONSE" | jq -r '.[].id' 2>/dev/null)
|
||||
MOVIE_COUNT=$(echo "$MOVIE_IDS" | grep -c "." 2>/dev/null || echo 0)
|
||||
|
||||
# Safety Layer 4 — movie count > 0
|
||||
if [[ "$MOVIE_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 movies — aborting to prevent mass deletion"
|
||||
notify "Radarr cleanup aborted on $(hostname) — 0 movies returned" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "Found $MOVIE_COUNT movies — fetching movie files..."
|
||||
|
||||
TMP_DIR="/tmp/radarr_cleanup_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "rm -rf $TMP_DIR" EXIT
|
||||
|
||||
TRACKED_FILE="$TMP_DIR/tracked_paths.txt"
|
||||
> "$TRACKED_FILE"
|
||||
|
||||
MOVIE_INDEX=0
|
||||
while IFS= read -r movie_id; do
|
||||
[[ -z "$movie_id" ]] && continue
|
||||
(( MOVIE_INDEX++ ))
|
||||
[[ $(( MOVIE_INDEX % 100 )) -eq 0 ]] && \
|
||||
log "Fetching files: $MOVIE_INDEX/$MOVIE_COUNT movies..."
|
||||
MOVIE_FILES=$(radarr_api "moviefile?movieId=${movie_id}" 2>/dev/null)
|
||||
if [[ -n "$MOVIE_FILES" ]]; then
|
||||
while IFS= read -r api_path; do
|
||||
[[ -z "$api_path" ]] && continue
|
||||
translate_path "$api_path" >> "$TRACKED_FILE"
|
||||
done < <(echo "$MOVIE_FILES" | jq -r '.[].path // .path' 2>/dev/null)
|
||||
fi
|
||||
done <<< "$MOVIE_IDS"
|
||||
|
||||
sort -u "$TRACKED_FILE" -o "$TRACKED_FILE"
|
||||
|
||||
# Build in-memory lookup map — O(1) per lookup vs O(n) grep per file
|
||||
# Eliminates the main performance bottleneck for large libraries
|
||||
declare -A TRACKED_MAP
|
||||
while IFS= read -r _tracked_path; do
|
||||
[[ -n "$_tracked_path" ]] && TRACKED_MAP["$_tracked_path"]=1
|
||||
done < "$TRACKED_FILE"
|
||||
unset _tracked_path
|
||||
info "Built in-memory lookup map: ${#TRACKED_MAP[@]} tracked paths"
|
||||
TRACKED_COUNT=$(wc -l < "$TRACKED_FILE")
|
||||
|
||||
# Safety Layer 5 — tracked count > 0
|
||||
if [[ "$TRACKED_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 tracked files — aborting to prevent mass deletion"
|
||||
notify "Radarr cleanup aborted on $(hostname) — 0 tracked files returned" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "$MOVIE_COUNT movies | $TRACKED_COUNT tracked movie files"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Scan Movies Root ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Scanning Movies Root ━━━"
|
||||
info "Root: $RADARR_MOVIES_ROOT | Orphan age: ${RADARR_ORPHAN_AGE} days"
|
||||
|
||||
START=$(date +%s)
|
||||
ORPHAN_COUNT=0
|
||||
JUNK_COUNT=0
|
||||
RECENT_COUNT=0
|
||||
PROTECTED_COUNT=0
|
||||
ORPHAN_BYTES=0
|
||||
JUNK_BYTES=0
|
||||
|
||||
AGE_SECONDS=$(( RADARR_ORPHAN_AGE * 86400 ))
|
||||
NOW=$(date +%s)
|
||||
MAX_DELETE_BYTES=$(awk "BEGIN {printf \"%d\", $RADARR_MAX_DELETE_GB * 1073741824}")
|
||||
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
|
||||
if [[ -n "${TRACKED_MAP[$filepath]:-}" ]]; then
|
||||
log "TRACKED: $filepath"
|
||||
continue
|
||||
fi
|
||||
|
||||
if is_protected_file "$filepath"; then
|
||||
log "$ICON_PROTECTED PROTECTED: $filepath"
|
||||
(( PROTECTED_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
FILE_SIZE=$(stat -c%s "$filepath" 2>/dev/null || echo 0)
|
||||
|
||||
if is_video_file "$filepath"; then
|
||||
FILE_MTIME=$(stat -c %Y "$filepath" 2>/dev/null || echo 0)
|
||||
FILE_AGE=$(( NOW - FILE_MTIME ))
|
||||
|
||||
if [[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && [[ "$SKIP_STRIKES" != true ]]; then
|
||||
log "RECENT (skipping): $filepath"
|
||||
(( RECENT_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
warn "$ICON_TRASH ORPHAN: $filepath"
|
||||
(( ORPHAN_COUNT++ ))
|
||||
ORPHAN_BYTES=$(( ORPHAN_BYTES + FILE_SIZE ))
|
||||
else
|
||||
log "JUNK: $filepath"
|
||||
(( JUNK_COUNT++ ))
|
||||
JUNK_BYTES=$(( JUNK_BYTES + FILE_SIZE ))
|
||||
fi
|
||||
|
||||
done < <(
|
||||
for host_path in "${ARR_PATH_MAP[@]}" "$RADARR_MOVIES_ROOT"; do
|
||||
[[ -d "$host_path" ]] && find "$host_path" -type f 2>/dev/null
|
||||
done | sort -u
|
||||
)
|
||||
|
||||
TOTAL_DELETE_BYTES=$(( ORPHAN_BYTES + JUNK_BYTES ))
|
||||
TOTAL_REMOVED=$(( ORPHAN_COUNT + JUNK_COUNT ))
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 6 — Deletion Size Threshold ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$TOTAL_DELETE_BYTES" -gt "$MAX_DELETE_BYTES" ]]; then
|
||||
TOTAL_HUMAN=$(awk "BEGIN {printf \"%.1fGB\", $TOTAL_DELETE_BYTES / 1073741824}")
|
||||
if [[ "$I_KNOW" != true ]]; then
|
||||
echo ""
|
||||
error "Deletion would exceed ${RADARR_MAX_DELETE_GB}GB — $TOTAL_HUMAN would be deleted"
|
||||
error "Review ORPHAN lines above carefully before proceeding"
|
||||
error "Rerun with: --i-know-what-im-doing"
|
||||
error "To also bypass age check: add --skip-strike-list"
|
||||
notify "Radarr cleanup halted on $(hostname) — ${TOTAL_HUMAN} requires --i-know-what-im-doing" \
|
||||
"Radarr Cleanup" "warning"
|
||||
exit 1
|
||||
else
|
||||
warn "OVERRIDE — deletion is $TOTAL_HUMAN — proceeding with --i-know-what-im-doing"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Execute Deletions ─────────────────────────────────────────────────────────────────────────
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
[[ -n "${TRACKED_MAP[$filepath]:-}" ]] && continue
|
||||
is_protected_file "$filepath" && continue
|
||||
|
||||
FILE_MTIME=$(stat -c %Y "$filepath" 2>/dev/null || echo 0)
|
||||
FILE_AGE=$(( NOW - FILE_MTIME ))
|
||||
|
||||
if is_video_file "$filepath"; then
|
||||
[[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && \
|
||||
[[ "$SKIP_STRIKES" != true ]] && continue
|
||||
fi
|
||||
|
||||
rm -f "$filepath" 2>/dev/null || error "Failed to delete: $filepath"
|
||||
|
||||
done < <(
|
||||
for host_path in "${ARR_PATH_MAP[@]}" "$RADARR_MOVIES_ROOT"; do
|
||||
[[ -d "$host_path" ]] && find "$host_path" -type f 2>/dev/null
|
||||
done | sort -u
|
||||
)
|
||||
|
||||
info "Cleaning up empty folders..."
|
||||
for host_path in "${ARR_PATH_MAP[@]}" "$RADARR_MOVIES_ROOT"; do
|
||||
[[ -d "$host_path" ]] && \
|
||||
find "$host_path" -mindepth 1 -type d -empty -delete 2>/dev/null
|
||||
done
|
||||
info "Empty folders removed"
|
||||
fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
ORPHAN_HUMAN=$(format_bytes "$ORPHAN_BYTES")
|
||||
JUNK_HUMAN=$(format_bytes "$JUNK_BYTES")
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY RADARR CLEANUP SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Tracked: $TRACKED_COUNT files ($MOVIE_COUNT movies)"
|
||||
echo "$ICON_SHIELD Protected: $PROTECTED_COUNT files (artwork, subtitles, metadata)"
|
||||
echo "$ICON_TRASH Orphans: $ORPHAN_COUNT files ($ORPHAN_HUMAN)"
|
||||
echo "$ICON_TRASH Junk: $JUNK_COUNT files ($JUNK_HUMAN)"
|
||||
echo "$ICON_SKIP Recent skipped: $RECENT_COUNT files (under ${RADARR_ORPHAN_AGE} days)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no files deleted"
|
||||
elif [[ "$TOTAL_REMOVED" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Clean — nothing to remove"
|
||||
else
|
||||
warn "$ICON_DONE Removed $TOTAL_REMOVED files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)"
|
||||
notify "Radarr cleanup on $(hostname) — removed $TOTAL_REMOVED files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)" \
|
||||
"Radarr Cleanup" "warning"
|
||||
# Notify Emby to clean missing files — removes ghost entries immediately
|
||||
notify_emby_scan
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# Write stats for sunday_morning_coffee_report.sh
|
||||
if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_CLEANUP_STATS:-}" ]]; then
|
||||
echo "$(date '+%Y-%m-%d')|radarr|${ORPHAN_COUNT}|${ORPHAN_BYTES}|${JUNK_COUNT}|${JUNK_BYTES}|${RECENT_COUNT}|${TRACKED_COUNT}" \
|
||||
>> "$ARR_CLEANUP_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
@@ -1,585 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Sonarr Cleanup =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Delete orphaned TV episode files not tracked by Sonarr. Queries the API for
|
||||
# all tracked episode file paths, walks the library on disk, and removes anything
|
||||
# untracked that is old enough to be past the import window. Triggers an Emby
|
||||
# library clean after each deletion run so ghost entries disappear immediately.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Every file encountered on disk is classified into one of five categories:
|
||||
#
|
||||
# TRACKED — Sonarr API knows this exact path → leave it alone
|
||||
# PROTECTED — matches SONARR_PROTECTED_PATTERNS → never delete
|
||||
# ORPHAN — video file, not tracked, older than SONARR_ORPHAN_AGE → delete
|
||||
# JUNK — not a video extension, not protected → delete regardless of age
|
||||
# RECENT — not tracked, under SONARR_ORPHAN_AGE → skip (may be mid-import)
|
||||
#
|
||||
# Sonarr generates show artwork (*.jpg), metadata (*.nfo), and manages subtitles
|
||||
# (*.srt, *.sub, *.ass) but does NOT include these in its tracked file API response.
|
||||
# Without PROTECTED classification these would be deleted — breaking Sonarr and
|
||||
# Emby metadata display.
|
||||
#
|
||||
# After deletions: notify_emby_scan() triggers Emby "Clean Missing Files" task.
|
||||
# Emby removes ghost entries immediately — no user-facing file-not-found errors.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Six gates — ALL must pass before any file is touched:
|
||||
# 1. Container running and not starting/unhealthy
|
||||
# 2. API reachable
|
||||
# 3. API version matches SONARR_VERSION_MAJOR in master.conf
|
||||
# 4. Series count > 0
|
||||
# 5. Tracked file count > 0
|
||||
# 6. Deletion size < SONARR_MAX_DELETE_GB — or --i-know-what-im-doing required
|
||||
#
|
||||
# acquire_lock "wait" — large scans take time, wait for previous run to finish
|
||||
# jq + curl validation — exits if either tool missing
|
||||
# DOCKER_TIMEOUT — container checks protected against daemon hangs
|
||||
# notify_emby_scan() — triggers Emby clean after deletion
|
||||
# validate_unraid_cmd — notify script validated before use
|
||||
# Silent by default — orphans/junk warn(), clean library logs silently
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_SONARR_URL / HOST*_SONARR_API_KEY / HOST*_SONARR_TV_ROOT
|
||||
# HOST*_SONARR_PATH_MAP — container path → host path translation
|
||||
# All aliased by detect_hosts() — script uses unprefixed names
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# SONARR_ORPHAN_AGE — days before untracked file eligible for deletion
|
||||
# SONARR_MAX_DELETE_GB — require --i-know-what-im-doing above this
|
||||
# SONARR_EXTENSIONS — video file extensions for orphan classification
|
||||
# SONARR_PROTECTED_PATTERNS — file patterns never deleted
|
||||
# SONARR_VERSION_MAJOR — expected Sonarr major version for API safety check
|
||||
# SONARR_IMPORT_SCAN_TIMEOUT — seconds to wait for pre-flight import scan (default 600)
|
||||
# ARR_CLEANUP_STATS — stats file path (read by coffee report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# sonarr_cleanup.sh — normal run
|
||||
# sonarr_cleanup.sh --dry-run — preview, no deletions
|
||||
# sonarr_cleanup.sh --log — verbose output
|
||||
# sonarr_cleanup.sh --status — show config and exit
|
||||
# sonarr_cleanup.sh --i-know-what-im-doing — bypass size threshold
|
||||
# sonarr_cleanup.sh --i-know-what-im-doing --skip-strike-list — NUCLEAR MODE
|
||||
#
|
||||
# NUCLEAR MODE: both flags bypass age check AND size threshold. User accepts full
|
||||
# responsibility — the flag name is long and annoying by design.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# ── Special flag pre-processing ───────────────────────────────────────────────────────────────
|
||||
I_KNOW=false
|
||||
SKIP_STRIKES=false
|
||||
FILTERED_ARGS=()
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--i-know-what-im-doing) I_KNOW=true ;;
|
||||
--skip-strike-list) SKIP_STRIKES=true ;;
|
||||
*) FILTERED_ARGS+=("$arg") ;;
|
||||
esac
|
||||
done
|
||||
|
||||
parse_args "${FILTERED_ARGS[@]}"
|
||||
|
||||
# ── Nuclear mode warning ──────────────────────────────────────────────────────────────────────
|
||||
if [[ "$I_KNOW" == true ]] && [[ "$SKIP_STRIKES" == true ]] && [[ "$DRY_RUN" != true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo "⚠️ WARNING — NUCLEAR MODE ACTIVE"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo " Flags: --i-know-what-im-doing --skip-strike-list"
|
||||
echo " Strike system: BYPASSED — deletes on first pass"
|
||||
echo " Size threshold: BYPASSED — no GB limit"
|
||||
echo " Data recovery: NOT POSSIBLE after deletion"
|
||||
echo ""
|
||||
echo " Review --dry-run output before proceeding."
|
||||
echo " You have 10 seconds to cancel (Ctrl+C)..."
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
sleep 10
|
||||
echo " Proceeding..."
|
||||
echo ""
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
error "curl not found — required for Sonarr API calls"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! command -v jq >/dev/null 2>&1; then
|
||||
error "jq not found — required for JSON parsing"
|
||||
notify "Sonarr cleanup failed on $(hostname) — jq not installed" "Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock "wait"
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases SONARR_URL, SONARR_API_KEY, SONARR_TV_ROOT
|
||||
detect_hosts
|
||||
|
||||
DOCKER_TIMEOUT=15
|
||||
SONARR_CONTAINER="Sonarr"
|
||||
|
||||
# Build path map from MY_ID's Sonarr path map
|
||||
declare -A ARR_PATH_MAP
|
||||
local_path_map_var="${MY_ID}_SONARR_PATH_MAP"
|
||||
eval "for key in \"\${!${local_path_map_var}[@]}\"; do
|
||||
ARR_PATH_MAP[\"\$key\"]=\"\${${local_path_map_var}[\$key]}\"
|
||||
done"
|
||||
|
||||
require_var SONARR_URL
|
||||
require_var SONARR_API_KEY
|
||||
require_var SONARR_TV_ROOT
|
||||
|
||||
if [[ ! -d "$SONARR_TV_ROOT" ]]; then
|
||||
error "TV root not found: $SONARR_TV_ROOT"
|
||||
notify "Sonarr cleanup failed on $(hostname) — TV root not found: $SONARR_TV_ROOT" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo " $MY_ID ($LOCAL_SERVER_NAME) — $SONARR_URL"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no files will be deleted"
|
||||
[[ "$I_KNOW" == true ]] && warn "OVERRIDE — --i-know-what-im-doing active"
|
||||
[[ "$SKIP_STRIKES" == true ]] && warn "OVERRIDE — --skip-strike-list active — age check bypassed"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY STATUS ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_GEAR Sonarr URL: $SONARR_URL"
|
||||
echo "$ICON_GEAR TV root: $SONARR_TV_ROOT"
|
||||
echo "$ICON_TIME Orphan age: ${SONARR_ORPHAN_AGE} days"
|
||||
echo "$ICON_GEAR Max delete: ${SONARR_MAX_DELETE_GB}GB (requires --i-know-what-im-doing)"
|
||||
echo "$ICON_GEAR Sonarr ver: v${SONARR_VERSION_MAJOR} expected"
|
||||
echo "$ICON_GEAR Extensions: ${SONARR_EXTENSIONS[*]}"
|
||||
echo "$ICON_GEAR Protected patterns: ${SONARR_PROTECTED_PATTERNS[*]}"
|
||||
echo "$ICON_GEAR Dry Run: $DRY_RUN"
|
||||
echo "$ICON_GEAR I know: $I_KNOW"
|
||||
echo "$ICON_GEAR Skip strikes: $SKIP_STRIKES"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 1 — Container Health ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SHIELD Safety Checks ━━━"
|
||||
|
||||
CONTAINER_RUNNING=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Running}}' "$SONARR_CONTAINER" 2>/dev/null)
|
||||
if [[ "$CONTAINER_RUNNING" != "true" ]]; then
|
||||
error "$SONARR_CONTAINER is not running — aborting"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — container not running" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
CONTAINER_HEALTH=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Health.Status}}' "$SONARR_CONTAINER" 2>/dev/null)
|
||||
case "$CONTAINER_HEALTH" in
|
||||
healthy) info "$SONARR_CONTAINER is healthy" ;;
|
||||
"") info "$SONARR_CONTAINER has no health check — proceeding" ;;
|
||||
starting)
|
||||
error "$SONARR_CONTAINER is still starting — aborting"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — container still starting" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1 ;;
|
||||
unhealthy)
|
||||
error "$SONARR_CONTAINER is unhealthy — aborting"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — container unhealthy" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1 ;;
|
||||
*) warn "$SONARR_CONTAINER health: $CONTAINER_HEALTH — proceeding with caution" ;;
|
||||
esac
|
||||
|
||||
info "Safety layer 1 passed — container healthy"
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
sonarr_api() {
|
||||
local endpoint="$1"
|
||||
local response http_code body
|
||||
|
||||
response=$(curl -sf \
|
||||
--max-time 30 \
|
||||
-H "X-Api-Key: $SONARR_API_KEY" \
|
||||
-w "\n%{http_code}" \
|
||||
"${SONARR_URL}/api/v3/${endpoint}" 2>/dev/null)
|
||||
|
||||
http_code=$(echo "$response" | tail -1)
|
||||
body=$(echo "$response" | head -n -1)
|
||||
|
||||
if [[ "$http_code" != "200" ]]; then
|
||||
error "Sonarr API HTTP $http_code for: $endpoint"
|
||||
return 1
|
||||
fi
|
||||
echo "$body"
|
||||
}
|
||||
|
||||
is_video_file() {
|
||||
local ext="${1##*.}"
|
||||
ext="${ext,,}"
|
||||
for valid_ext in "${SONARR_EXTENSIONS[@]}"; do
|
||||
[[ "$ext" == "$valid_ext" ]] && return 0
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
is_protected_file() {
|
||||
local filename
|
||||
filename=$(basename "$1")
|
||||
for pattern in "${SONARR_PROTECTED_PATTERNS[@]}"; do
|
||||
# shellcheck disable=SC2254
|
||||
case "$filename" in
|
||||
$pattern) return 0 ;;
|
||||
esac
|
||||
done
|
||||
return 1
|
||||
}
|
||||
|
||||
format_bytes() {
|
||||
local bytes=$1
|
||||
if (( bytes > 1073741824 )); then
|
||||
awk "BEGIN {printf \"%.1fGB\", $bytes / 1073741824}"
|
||||
elif (( bytes > 1048576 )); then
|
||||
awk "BEGIN {printf \"%.1fMB\", $bytes / 1048576}"
|
||||
else
|
||||
echo "${bytes}B"
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Pre-flight: Sonarr Import Scan ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Pre-flight: Sonarr Import Scan ━━━"
|
||||
|
||||
# Reverse-lookup container path from path map so Sonarr gets its own path, not the host path
|
||||
SONARR_CONTAINER_ROOT=""
|
||||
for _cp in "${!ARR_PATH_MAP[@]}"; do
|
||||
if [[ "${ARR_PATH_MAP[$_cp]}" == "$SONARR_TV_ROOT" ]]; then
|
||||
SONARR_CONTAINER_ROOT="$_cp"
|
||||
break
|
||||
fi
|
||||
done
|
||||
unset _cp
|
||||
|
||||
if [[ -n "$SONARR_CONTAINER_ROOT" ]]; then
|
||||
info "Triggering DownloadedEpisodesScan on: $SONARR_CONTAINER_ROOT"
|
||||
SCAN_PAYLOAD="{\"name\": \"DownloadedEpisodesScan\", \"path\": \"$SONARR_CONTAINER_ROOT\"}"
|
||||
else
|
||||
info "No path map match — triggering DownloadedEpisodesScan (all root folders)"
|
||||
SCAN_PAYLOAD='{"name": "DownloadedEpisodesScan"}'
|
||||
fi
|
||||
|
||||
SCAN_RESPONSE=$(curl -sf --max-time 30 -X POST \
|
||||
-H "X-Api-Key: $SONARR_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$SCAN_PAYLOAD" \
|
||||
"${SONARR_URL}/api/v3/command" 2>/dev/null)
|
||||
|
||||
SCAN_CMD_ID=$(echo "$SCAN_RESPONSE" | jq -r '.id // empty' 2>/dev/null)
|
||||
|
||||
if [[ -z "$SCAN_CMD_ID" ]]; then
|
||||
warn "Could not trigger import scan — proceeding without pre-flight"
|
||||
else
|
||||
info "Import scan queued (command ID: $SCAN_CMD_ID) — waiting for completion..."
|
||||
POLL_TIMEOUT=${SONARR_IMPORT_SCAN_TIMEOUT:-600}
|
||||
POLLED=0
|
||||
while [[ "$POLLED" -lt "$POLL_TIMEOUT" ]]; do
|
||||
SCAN_STATUS=$(curl -sf --max-time 10 \
|
||||
-H "X-Api-Key: $SONARR_API_KEY" \
|
||||
"${SONARR_URL}/api/v3/command/${SCAN_CMD_ID}" 2>/dev/null | \
|
||||
jq -r '.status // empty' 2>/dev/null)
|
||||
case "$SCAN_STATUS" in
|
||||
completed) info "Import scan complete ✅"; break ;;
|
||||
failed) warn "Import scan reported failed — proceeding anyway"; break ;;
|
||||
esac
|
||||
sleep 10
|
||||
(( POLLED += 10 ))
|
||||
[[ $(( POLLED % 60 )) -eq 0 ]] && log " Still scanning... (${POLLED}s elapsed)"
|
||||
done
|
||||
[[ "$POLLED" -ge "$POLL_TIMEOUT" ]] && \
|
||||
warn "Import scan timed out after ${POLL_TIMEOUT}s — proceeding anyway"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Fetch Sonarr Tracked Files ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Fetching Sonarr Tracked Files ━━━"
|
||||
|
||||
# Safety Layer 2 — API reachability
|
||||
if ! check_api "$SONARR_URL" "Sonarr" 10; then
|
||||
notify "Sonarr cleanup aborted on $(hostname) — API unreachable" "Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Safety Layer 3 — API version check
|
||||
check_arr_version "$SONARR_URL" "$SONARR_API_KEY" "v3" "$SONARR_VERSION_MAJOR" "Sonarr" || exit 1
|
||||
|
||||
info "Querying Sonarr API..."
|
||||
|
||||
# Fetch all series
|
||||
SERIES_RESPONSE=$(sonarr_api "series") || {
|
||||
error "Failed to fetch series from Sonarr"
|
||||
notify "Sonarr cleanup failed on $(hostname) — could not fetch series" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
}
|
||||
|
||||
SERIES_IDS=$(echo "$SERIES_RESPONSE" | jq -r '.[].id' 2>/dev/null)
|
||||
SERIES_COUNT=$(echo "$SERIES_IDS" | grep -c "." 2>/dev/null || echo 0)
|
||||
|
||||
# Safety Layer 4 — series count > 0
|
||||
if [[ "$SERIES_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 series — aborting to prevent mass deletion"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — 0 series returned" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "Found $SERIES_COUNT series — fetching episode files..."
|
||||
|
||||
TMP_DIR="/tmp/sonarr_cleanup_$$"
|
||||
mkdir -p "$TMP_DIR"
|
||||
trap "rm -rf $TMP_DIR" EXIT
|
||||
|
||||
TRACKED_FILE="$TMP_DIR/tracked_paths.txt"
|
||||
> "$TRACKED_FILE"
|
||||
|
||||
SERIES_INDEX=0
|
||||
while IFS= read -r series_id; do
|
||||
[[ -z "$series_id" ]] && continue
|
||||
(( SERIES_INDEX++ ))
|
||||
[[ $(( SERIES_INDEX % 50 )) -eq 0 ]] && \
|
||||
log "Fetching files: $SERIES_INDEX/$SERIES_COUNT series..."
|
||||
SERIES_FILES=$(sonarr_api "episodefile?seriesId=${series_id}" 2>/dev/null)
|
||||
if [[ -n "$SERIES_FILES" ]]; then
|
||||
while IFS= read -r api_path; do
|
||||
[[ -z "$api_path" ]] && continue
|
||||
translate_path "$api_path" >> "$TRACKED_FILE"
|
||||
done < <(echo "$SERIES_FILES" | jq -r '.[].path' 2>/dev/null)
|
||||
fi
|
||||
done <<< "$SERIES_IDS"
|
||||
|
||||
sort -u "$TRACKED_FILE" -o "$TRACKED_FILE"
|
||||
|
||||
# Build in-memory lookup map — O(1) per lookup vs O(n) grep per file
|
||||
declare -A TRACKED_MAP
|
||||
while IFS= read -r _tracked_path; do
|
||||
[[ -n "$_tracked_path" ]] && TRACKED_MAP["$_tracked_path"]=1
|
||||
done < "$TRACKED_FILE"
|
||||
unset _tracked_path
|
||||
info "Built in-memory lookup map: ${#TRACKED_MAP[@]} tracked paths"
|
||||
TRACKED_COUNT=$(wc -l < "$TRACKED_FILE")
|
||||
|
||||
# Safety Layer 5 — tracked count > 0
|
||||
if [[ "$TRACKED_COUNT" -eq 0 ]]; then
|
||||
error "API returned 0 tracked files — aborting to prevent mass deletion"
|
||||
notify "Sonarr cleanup aborted on $(hostname) — 0 tracked files returned" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
info "$SERIES_COUNT series | $TRACKED_COUNT tracked episode files"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Scan TV Root ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_CLEAN Scanning TV Root ━━━"
|
||||
info "Root: $SONARR_TV_ROOT | Orphan age: ${SONARR_ORPHAN_AGE} days"
|
||||
|
||||
START=$(date +%s)
|
||||
ORPHAN_COUNT=0
|
||||
JUNK_COUNT=0
|
||||
RECENT_COUNT=0
|
||||
PROTECTED_COUNT=0
|
||||
ORPHAN_BYTES=0
|
||||
JUNK_BYTES=0
|
||||
|
||||
AGE_SECONDS=$(( SONARR_ORPHAN_AGE * 86400 ))
|
||||
NOW=$(date +%s)
|
||||
MAX_DELETE_BYTES=$(awk "BEGIN {printf \"%d\", $SONARR_MAX_DELETE_GB * 1073741824}")
|
||||
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
|
||||
if [[ -n "${TRACKED_MAP[$filepath]:-}" ]]; then
|
||||
log "TRACKED: $filepath"
|
||||
continue
|
||||
fi
|
||||
|
||||
if is_protected_file "$filepath"; then
|
||||
log "$ICON_PROTECTED PROTECTED: $filepath"
|
||||
(( PROTECTED_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
FILE_SIZE=$(stat -c%s "$filepath" 2>/dev/null || echo 0)
|
||||
|
||||
if is_video_file "$filepath"; then
|
||||
FILE_MTIME=$(stat -c %Y "$filepath" 2>/dev/null || echo 0)
|
||||
FILE_AGE=$(( NOW - FILE_MTIME ))
|
||||
|
||||
if [[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && [[ "$SKIP_STRIKES" != true ]]; then
|
||||
log "RECENT (skipping): $filepath"
|
||||
(( RECENT_COUNT++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
warn "$ICON_TRASH ORPHAN: $filepath"
|
||||
(( ORPHAN_COUNT++ ))
|
||||
ORPHAN_BYTES=$(( ORPHAN_BYTES + FILE_SIZE ))
|
||||
else
|
||||
log "JUNK: $filepath"
|
||||
(( JUNK_COUNT++ ))
|
||||
JUNK_BYTES=$(( JUNK_BYTES + FILE_SIZE ))
|
||||
fi
|
||||
|
||||
done < <(
|
||||
for host_path in "${ARR_PATH_MAP[@]}" "$SONARR_TV_ROOT"; do
|
||||
[[ -d "$host_path" ]] && find "$host_path" -type f 2>/dev/null
|
||||
done | sort -u
|
||||
)
|
||||
|
||||
TOTAL_DELETE_BYTES=$(( ORPHAN_BYTES + JUNK_BYTES ))
|
||||
TOTAL_REMOVED=$(( ORPHAN_COUNT + JUNK_COUNT ))
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Safety Layer 6 — Deletion Size Threshold ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$TOTAL_DELETE_BYTES" -gt "$MAX_DELETE_BYTES" ]]; then
|
||||
TOTAL_HUMAN=$(awk "BEGIN {printf \"%.1fGB\", $TOTAL_DELETE_BYTES / 1073741824}")
|
||||
if [[ "$I_KNOW" != true ]]; then
|
||||
echo ""
|
||||
error "Deletion would exceed ${SONARR_MAX_DELETE_GB}GB — $TOTAL_HUMAN would be deleted"
|
||||
error "Review ORPHAN lines above carefully before proceeding"
|
||||
error "Rerun with: --i-know-what-im-doing"
|
||||
error "To also bypass age check: add --skip-strike-list"
|
||||
notify "Sonarr cleanup halted on $(hostname) — ${TOTAL_HUMAN} requires --i-know-what-im-doing" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
exit 1
|
||||
else
|
||||
warn "OVERRIDE — deletion is $TOTAL_HUMAN — proceeding with --i-know-what-im-doing"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ── Execute Deletions ─────────────────────────────────────────────────────────────────────────
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
while IFS= read -r filepath; do
|
||||
[[ -z "$filepath" ]] && continue
|
||||
[[ -n "${TRACKED_MAP[$filepath]:-}" ]] && continue
|
||||
is_protected_file "$filepath" && continue
|
||||
|
||||
FILE_MTIME=$(stat -c %Y "$filepath" 2>/dev/null || echo 0)
|
||||
FILE_AGE=$(( NOW - FILE_MTIME ))
|
||||
|
||||
if is_video_file "$filepath"; then
|
||||
[[ "$FILE_AGE" -lt "$AGE_SECONDS" ]] && \
|
||||
[[ "$SKIP_STRIKES" != true ]] && continue
|
||||
fi
|
||||
|
||||
rm -f "$filepath" 2>/dev/null || error "Failed to delete: $filepath"
|
||||
|
||||
done < <(
|
||||
for host_path in "${ARR_PATH_MAP[@]}" "$SONARR_TV_ROOT"; do
|
||||
[[ -d "$host_path" ]] && find "$host_path" -type f 2>/dev/null
|
||||
done | sort -u
|
||||
)
|
||||
|
||||
info "Cleaning up empty folders..."
|
||||
for host_path in "${ARR_PATH_MAP[@]}" "$SONARR_TV_ROOT"; do
|
||||
[[ -d "$host_path" ]] && \
|
||||
find "$host_path" -mindepth 1 -type d -empty -delete 2>/dev/null
|
||||
done
|
||||
info "Empty folders removed"
|
||||
fi
|
||||
|
||||
END=$(date +%s)
|
||||
|
||||
ORPHAN_HUMAN=$(format_bytes "$ORPHAN_BYTES")
|
||||
JUNK_HUMAN=$(format_bytes "$JUNK_BYTES")
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY SONARR CLEANUP SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SYNC Tracked: $TRACKED_COUNT files ($SERIES_COUNT series)"
|
||||
echo "$ICON_SHIELD Protected: $PROTECTED_COUNT files (artwork, subtitles, metadata)"
|
||||
echo "$ICON_TRASH Orphans: $ORPHAN_COUNT files ($ORPHAN_HUMAN)"
|
||||
echo "$ICON_TRASH Junk: $JUNK_COUNT files ($JUNK_HUMAN)"
|
||||
echo "$ICON_SKIP Recent skipped: $RECENT_COUNT files (under ${SONARR_ORPHAN_AGE} days)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no files deleted"
|
||||
elif [[ "$TOTAL_REMOVED" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Clean — nothing to remove"
|
||||
else
|
||||
warn "$ICON_DONE Removed $TOTAL_REMOVED files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)"
|
||||
notify "Sonarr cleanup on $(hostname) — removed $TOTAL_REMOVED files (orphans: $ORPHAN_HUMAN junk: $JUNK_HUMAN)" \
|
||||
"Sonarr Cleanup" "warning"
|
||||
# Notify Emby to clean missing files — removes ghost entries immediately
|
||||
notify_emby_scan
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# Write stats for sunday_morning_coffee_report.sh
|
||||
if [[ "$DRY_RUN" == false ]] && [[ -n "${ARR_CLEANUP_STATS:-}" ]]; then
|
||||
echo "$(date '+%Y-%m-%d')|sonarr|${ORPHAN_COUNT}|${ORPHAN_BYTES}|${JUNK_COUNT}|${JUNK_BYTES}|${RECENT_COUNT}|${TRACKED_COUNT}" \
|
||||
>> "$ARR_CLEANUP_STATS" 2>/dev/null || true
|
||||
fi
|
||||
|
||||
exit 0
|
||||
+10
-10
@@ -151,14 +151,14 @@ BACKUP_VERIFY_MIN_SIZE="1M" # skip files smaller than this
|
||||
|
||||
```bash
|
||||
# master.conf
|
||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db" # survives reboots
|
||||
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db" # survives reboots
|
||||
BANDWIDTH_LOG_RETENTION=90 # days — file stays bounded, never grows unbounded
|
||||
BANDWIDTH_WARN_GB=50 # flag transfers or daily totals exceeding this
|
||||
```
|
||||
|
||||
**Why `/boot/config/`**: The log needs to survive reboots to build a useful history.
|
||||
`/boot/config/` is on the USB flash drive, which survives reboots and is backed up
|
||||
by unRAID's flash backup. The log is bounded by `BANDWIDTH_LOG_RETENTION` so it never
|
||||
**Why `$DATA_DIR`**: The log needs to survive reboots to build a useful history.
|
||||
`$DATA_DIR` (`$SCRIPTS_DIR/data/`) is on the boot device (internal) or appdata (flash),
|
||||
both of which survive reboots. The log is bounded by `BANDWIDTH_LOG_RETENTION` so it never
|
||||
grows unbounded.
|
||||
|
||||
**`BANDWIDTH_WARN_GB`**: Set to a value that represents "unexpectedly large" for your
|
||||
@@ -230,7 +230,7 @@ data — it just doesn't trigger a notification for that condition.
|
||||
| State File | Source | What It Shows |
|
||||
|-----------|--------|---------------|
|
||||
| `FALLBACK_STATE_FILE` | `Fallback/fallback.sh` | Current fallback state (NORMAL/FALLBACK/etc.) |
|
||||
| `SYS_WATCHDOG_FAILED_FILE` | `Watchdogs/docker_watchdog.sh` | Container skip list — needs human attention |
|
||||
| `DOCKER_WATCHDOG_FAILED_FILE` | `Watchdogs/docker_watchdog.sh` | Container skip list — needs human attention |
|
||||
| `WATCHDOG_STATE_FILE` | `Watchdogs/docker_watchdog.sh` | Active container strike counts |
|
||||
| `SYS_WATCHDOG_STATE_FILE` | `Watchdogs/stability_watchdog.sh` | Active system watchdog strikes |
|
||||
| `BANDWIDTH_LOG` | `bandwidth_monitor.sh` | Yesterday's transfer history |
|
||||
@@ -254,7 +254,7 @@ HOST2_EMBY_API_KEY="<host2_api_key>"
|
||||
```
|
||||
|
||||
To generate an API key: Emby UI → Settings → API Keys → New API Key.
|
||||
Give it a descriptive name (e.g., `unraid_scripts`). The key is only shown once.
|
||||
Give it a descriptive name (e.g., `varaverk`). The key is only shown once.
|
||||
|
||||
### Report Configuration
|
||||
|
||||
@@ -285,7 +285,7 @@ higher than 90% — at 100% utilisation new inotify watches silently fail.
|
||||
**`PHP_FPM_WARN_PCT`**: 80% means the WebGUI is using most of its workers. At 100%
|
||||
new requests queue (WebGUI feels sluggish) or time out.
|
||||
|
||||
**`PHP_MAX_CHILDREN`**: Set by `php_fpm_max_children.sh` in `unRAID_Essentials/` — do
|
||||
**`PHP_MAX_CHILDREN`**: Set by `php_fpm_max_children.sh` in `System_Essentials/` — do
|
||||
not set manually here.
|
||||
|
||||
### Reading the Weekly Digest Data
|
||||
@@ -297,7 +297,7 @@ shows for the week:
|
||||
|
||||
A few warning snapshots per week is normal. A rising peak or many warnings per week
|
||||
means the limits should be adjusted — use `inotify_tuning.sh` or `php_fpm_max_children.sh`
|
||||
in `unRAID_Essentials/`.
|
||||
in `System_Essentials/`.
|
||||
|
||||
---
|
||||
|
||||
@@ -401,7 +401,7 @@ BACKUP_VERIFY_SAMPLE=10
|
||||
BACKUP_VERIFY_MIN_SIZE="1M"
|
||||
|
||||
# bandwidth_monitor.sh
|
||||
BANDWIDTH_LOG="/boot/config/bandwidth_history.db"
|
||||
BANDWIDTH_LOG="$DATA_DIR/bandwidth_history.db"
|
||||
BANDWIDTH_LOG_RETENTION=90
|
||||
BANDWIDTH_WARN_GB=50
|
||||
|
||||
@@ -566,7 +566,7 @@ SILENT_MODE=false # monitor script — output is the point
|
||||
parse_args "$@"
|
||||
|
||||
# root check if needed
|
||||
# tool validation (validate_unraid_cmd)
|
||||
# tool validation (platform_require_cmd)
|
||||
acquire_lock
|
||||
detect_hosts # if host-specific config needed
|
||||
|
||||
|
||||
@@ -82,13 +82,13 @@ drive SMART attributes, ZFS pool state, ARC statistics, kernel memory pressure.
|
||||
```
|
||||
Monitors/ ← observes and reports (this folder)
|
||||
Docker_Essentials/ ← acts on containers (docker_watchdog starts/stops)
|
||||
unRAID_Essentials/ ← acts on the server (stability_watchdog, inotify_tuning)
|
||||
Fallback/ ← acts on the full stack (failover, handback)
|
||||
System_Essentials/ ← acts on the server (stability_watchdog, inotify_tuning)
|
||||
Fallback/ ← acts on the full stack (fallback, handback)
|
||||
Rsync/ ← calls bandwidth_monitor.sh (auto-logs each sync)
|
||||
```
|
||||
|
||||
`weekly_health_digest.sh` reads state files written by scripts in Docker_Essentials,
|
||||
unRAID_Essentials, Fallback, and Rsync. It is the only script in this folder with
|
||||
System_Essentials, Fallback, and Rsync. It is the only script in this folder with
|
||||
runtime dependencies on other folders' output — everything else is fully independent.
|
||||
|
||||
---
|
||||
@@ -153,7 +153,7 @@ Daily 8am:
|
||||
weekly_health_digest.sh ── reads ──────────► FALLBACK_STATE_FILE
|
||||
── reads ──────────► WATCHDOG_STATE_FILE
|
||||
── reads ──────────► SYS_WATCHDOG_STATE_FILE
|
||||
── reads ──────────► SYS_WATCHDOG_FAILED_FILE
|
||||
── reads ──────────► DOCKER_WATCHDOG_FAILED_FILE
|
||||
── reads ──────────► BANDWIDTH_LOG
|
||||
── reads ──────────► TRANSCODE_DAILY_LOG
|
||||
── reads ──────────► TUNING_MONITOR_LOG
|
||||
|
||||
+74
-19
@@ -19,6 +19,24 @@
|
||||
# to HOST*_DAILY_SYNC_SHARES. Both aliased by detect_hosts().
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Pre-flight — connectivity to the remote, remote array mounted, version parity
|
||||
# 2. Resolve the share list (BACKUP_VERIFY_SHARES, else DAILY_SYNC_SHARES)
|
||||
# 3. Per share:
|
||||
# a. Randomly sample BACKUP_VERIFY_SAMPLE files above BACKUP_VERIFY_MIN_SIZE
|
||||
# b. Compute each file's MD5 locally
|
||||
# c. Compute the same file's MD5 on the remote over SSH
|
||||
# d. Classify: MATCH | MISMATCH | MISSING
|
||||
# 4. Report per-share and overall counts; notify on any MISMATCH and on
|
||||
# significant MISSING counts
|
||||
#
|
||||
# Sampling rather than full verification is deliberate — a complete checksum of every
|
||||
# mirrored file would take longer than the interval between runs. Random sampling over
|
||||
# a weekly cadence surfaces systematic corruption without ever reading the whole library.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -56,9 +74,6 @@
|
||||
# SSH Timeout
|
||||
# SSH_TIMEOUT caps all SSH calls. One hung connection does not block the run.
|
||||
#
|
||||
# Notification Validated
|
||||
# validate_unraid_cmd confirms the notify script is present before use.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
@@ -107,6 +122,7 @@ source "$SCRIPT_DIR/../load_config.sh"
|
||||
parse_args "$@"
|
||||
|
||||
SSH_TIMEOUT=15
|
||||
BACKUP_VERIFY_MD5_TIMEOUT_MAX="${BACKUP_VERIFY_MD5_TIMEOUT_MAX:-600}"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
@@ -119,15 +135,12 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases BACKUP_VERIFY_SHARES + DAILY_SYNC_SHARES
|
||||
detect_hosts
|
||||
require_partnership
|
||||
|
||||
# Share selection — configured list or fallback to daily sync shares
|
||||
if [[ ${#BACKUP_VERIFY_SHARES[@]} -gt 0 ]]; then
|
||||
@@ -144,6 +157,10 @@ if [[ ${#VERIFY_SHARES[@]} -eq 0 ]]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Config: sample=${BACKUP_VERIFY_SAMPLE} min-size=${BACKUP_VERIFY_MIN_SIZE} ssh-timeout=${SSH_TIMEOUT}s"
|
||||
log "$ICON_GEAR Remote: $REMOTE_ID ($REMOTE_SERVER_NAME — $REMOTE_SERVER)"
|
||||
log "$ICON_GEAR Shares: ${VERIFY_SHARES[*]}"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — showing sample selection only, no checksums computed"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -177,12 +194,14 @@ resolve_remote_ip
|
||||
|
||||
# Connectivity — no point making 100+ SSH calls if remote is unreachable
|
||||
check_connectivity
|
||||
echo "Connectivity to $REMOTE_SERVER_NAME ✅"
|
||||
|
||||
# Version parity — mismatched unRAID could cause md5sum path differences
|
||||
check_unraid_version_parity || {
|
||||
check_os_version_parity || {
|
||||
warn "Version parity check failed — proceeding with caution"
|
||||
warn "Checksum results may be unreliable if md5sum path changed between versions"
|
||||
}
|
||||
echo "Version parity with $REMOTE_SERVER_NAME ✅"
|
||||
|
||||
# Remote array — if array is down all files appear "missing" = false alarm
|
||||
if ! check_remote_array; then
|
||||
@@ -192,6 +211,7 @@ if ! check_remote_array; then
|
||||
"Backup Verify" "warning"
|
||||
exit 1
|
||||
fi
|
||||
echo "Remote array mounted on $REMOTE_SERVER_NAME ✅"
|
||||
|
||||
echo "Pre-flight passed ✅"
|
||||
|
||||
@@ -209,6 +229,7 @@ TOTAL_CHECKED=0
|
||||
TOTAL_MATCH=0
|
||||
TOTAL_MISMATCH=0
|
||||
TOTAL_MISSING=0
|
||||
TOTAL_UNVERIFIED=0
|
||||
SHARES_WITH_ISSUES=()
|
||||
|
||||
for share in "${VERIFY_SHARES[@]}"; do
|
||||
@@ -246,6 +267,7 @@ for share in "${VERIFY_SHARES[@]}"; do
|
||||
SHARE_MATCH=0
|
||||
SHARE_MISMATCH=0
|
||||
SHARE_MISSING=0
|
||||
SHARE_UNVERIFIED=0
|
||||
|
||||
for local_file in "${SAMPLE_FILES[@]}"; do
|
||||
[[ -z "$local_file" ]] && continue
|
||||
@@ -257,19 +279,49 @@ for share in "${VERIFY_SHARES[@]}"; do
|
||||
continue
|
||||
fi
|
||||
|
||||
# Remote checksum via SSH — timeout protected
|
||||
remote_md5=$(timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
|
||||
# The path is interpolated into a remote shell command, so it must be escaped for
|
||||
# reuse as one word. A bare '$local_file' inside single quotes breaks on the first
|
||||
# apostrophe — "Frieren - Beyond Journey's End" ended the quote early, md5sum fell
|
||||
# back to reading stdin, and the empty-input hash d41d8cd9... was reported as a
|
||||
# MISMATCH against a file that is byte-identical on the remote.
|
||||
printf -v remote_q '%q' "$local_file"
|
||||
|
||||
# Existence and content are separate questions. Asking them together means a slow
|
||||
# checksum is indistinguishable from an absent file.
|
||||
remote_exists=$(timeout "$SSH_TIMEOUT" ssh -i "$SSH_KEY" \
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o StrictHostKeyChecking=no \
|
||||
root@"$REMOTE_SERVER" \
|
||||
"md5sum '$local_file' 2>/dev/null | awk '{print \$1}'" 2>/dev/null)
|
||||
"test -f $remote_q && echo yes" 2>/dev/null </dev/null)
|
||||
|
||||
(( TOTAL_CHECKED++ ))
|
||||
|
||||
if [[ -z "$remote_md5" ]]; then
|
||||
if [[ "$remote_exists" != "yes" ]]; then
|
||||
warn "$ICON_ERROR MISSING: $(basename "$local_file")"
|
||||
(( SHARE_MISSING++ ))
|
||||
(( TOTAL_MISSING++ ))
|
||||
continue
|
||||
fi
|
||||
|
||||
# md5sum of a multi-GB file cannot finish inside a connect-sized timeout. Budget by
|
||||
# size — a 5.9GB file needs ~30s and was being killed at 15s, then counted MISSING
|
||||
# even though it was present and correct.
|
||||
local_size=$(stat -c%s "$local_file" 2>/dev/null || echo 0)
|
||||
md5_timeout=$(( local_size / 52428800 + SSH_TIMEOUT ))
|
||||
(( md5_timeout > BACKUP_VERIFY_MD5_TIMEOUT_MAX )) && md5_timeout=$BACKUP_VERIFY_MD5_TIMEOUT_MAX
|
||||
|
||||
remote_md5=$(timeout "$md5_timeout" ssh -i "$SSH_KEY" \
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
-o StrictHostKeyChecking=no \
|
||||
root@"$REMOTE_SERVER" \
|
||||
"md5sum $remote_q 2>/dev/null | awk '{print \$1}'" 2>/dev/null </dev/null)
|
||||
|
||||
if [[ -z "$remote_md5" ]]; then
|
||||
# Present but unreadable within budget. Reporting this as a mismatch or a miss
|
||||
# would be a claim the run did not earn.
|
||||
warn "$ICON_WARN UNVERIFIED (checksum timed out after ${md5_timeout}s): $(basename "$local_file")"
|
||||
(( SHARE_UNVERIFIED++ ))
|
||||
(( TOTAL_UNVERIFIED++ ))
|
||||
elif [[ "$local_md5" == "$remote_md5" ]]; then
|
||||
log "MATCH: $(basename "$local_file")"
|
||||
(( SHARE_MATCH++ ))
|
||||
@@ -284,11 +336,11 @@ for share in "${VERIFY_SHARES[@]}"; do
|
||||
done
|
||||
|
||||
# Per-share result — only visible if issues found
|
||||
if [[ "$SHARE_MISMATCH" -gt 0 || "$SHARE_MISSING" -gt 0 ]]; then
|
||||
warn "$SHARE_NAME — match: $SHARE_MATCH missing: $SHARE_MISSING mismatch: $SHARE_MISMATCH"
|
||||
if [[ "$SHARE_MISMATCH" -gt 0 || "$SHARE_MISSING" -gt 0 || "$SHARE_UNVERIFIED" -gt 0 ]]; then
|
||||
warn "$SHARE_NAME — match: $SHARE_MATCH missing: $SHARE_MISSING mismatch: $SHARE_MISMATCH unverified: $SHARE_UNVERIFIED"
|
||||
SHARES_WITH_ISSUES+=("$SHARE_NAME")
|
||||
else
|
||||
log "$SHARE_NAME — all $SHARE_MATCH files match ✅"
|
||||
echo "$SHARE_NAME — all $SHARE_MATCH files match ✅"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
@@ -306,10 +358,11 @@ echo "$ICON_VERIFY Checked: $TOTAL_CHECKED files"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$TOTAL_MISMATCH" -gt 0 || "$TOTAL_MISSING" -gt 0 ]]; then
|
||||
echo "$ICON_SUCCESS Match: $TOTAL_MATCH"
|
||||
warn "Missing: $TOTAL_MISSING"
|
||||
[[ "$TOTAL_MISMATCH" -gt 0 ]] && echo "$ICON_ERROR Mismatch: $TOTAL_MISMATCH"
|
||||
if [[ "$TOTAL_MISMATCH" -gt 0 || "$TOTAL_MISSING" -gt 0 || "$TOTAL_UNVERIFIED" -gt 0 ]]; then
|
||||
echo "$ICON_SUCCESS Match: $TOTAL_MATCH"
|
||||
warn "Missing: $TOTAL_MISSING"
|
||||
[[ "$TOTAL_MISMATCH" -gt 0 ]] && echo "$ICON_ERROR Mismatch: $TOTAL_MISMATCH"
|
||||
[[ "$TOTAL_UNVERIFIED" -gt 0 ]] && warn "Unverified: $TOTAL_UNVERIFIED (present, checksum timed out)"
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
@@ -318,6 +371,8 @@ elif [[ "$TOTAL_MISMATCH" -gt 0 || "$TOTAL_MISSING" -gt 0 ]]; then
|
||||
echo "$ICON_ERROR Status: ISSUES FOUND — ${#SHARES_WITH_ISSUES[@]} share(s) need attention: ${SHARES_WITH_ISSUES[*]}"
|
||||
notify "Backup verify FAILED on $(hostname) → $REMOTE_SERVER_NAME — mismatches: $TOTAL_MISMATCH missing: $TOTAL_MISSING — shares: ${SHARES_WITH_ISSUES[*]}" \
|
||||
"Backup Verify" "warning"
|
||||
elif [[ "$TOTAL_UNVERIFIED" -gt 0 ]]; then
|
||||
warn "Status: $TOTAL_MATCH verified, $TOTAL_UNVERIFIED could not be checksummed in time — NOT a clean run"
|
||||
else
|
||||
echo "$ICON_DONE Status: all $TOTAL_CHECKED files match across ${#VERIFY_SHARES[@]} shares ✅"
|
||||
fi
|
||||
|
||||
Regular → Executable
+20
-7
@@ -17,6 +17,24 @@
|
||||
# last 7 days activity timeline, any transfers or days exceeding BANDWIDTH_WARN_GB.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Log mode (--log-transfer), called by rsync.sh after every sync:
|
||||
# 1. Append one line: YYYY-MM-DD|HH:MM|profile|duration|status|bytes
|
||||
# 2. Trim entries older than BANDWIDTH_LOG_RETENTION days
|
||||
# One bounded write per rsync run — never grows without limit, never rewrites history.
|
||||
#
|
||||
# Report mode (default), scheduled weekly:
|
||||
# 1. Read the accumulated log
|
||||
# 2. Aggregate per profile — run count, total bytes, average duration, failures
|
||||
# 3. Build a 7-day activity timeline
|
||||
# 4. Flag any single transfer or any single day exceeding BANDWIDTH_WARN_GB
|
||||
#
|
||||
# The two modes never run together: logging is a side effect of rsync, reporting is a
|
||||
# scheduled read. Report mode never writes to the log.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -46,9 +64,6 @@
|
||||
# Log Directory Guard
|
||||
# Creates the log directory if it doesn't exist. Exits cleanly if unwritable.
|
||||
#
|
||||
# Notification Validated
|
||||
# validate_unraid_cmd confirms the notify script is present before use.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
@@ -134,14 +149,12 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
# detect_hosts() sets MY_ID for report header
|
||||
detect_hosts
|
||||
|
||||
log "$ICON_GEAR Config: log=${BANDWIDTH_LOG} retention=${BANDWIDTH_LOG_RETENTION}d warn=${BANDWIDTH_WARN_GB}GB"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
|
||||
Regular → Executable
+64
-39
@@ -15,6 +15,23 @@
|
||||
# separate message lists all CRITICAL domains. Not one notification per domain.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Validate openssl is present (platform_require_cmd) — without it nothing can be checked
|
||||
# 2. Per domain in CERT_MONITOR_DOMAINS:
|
||||
# a. Open a real TLS connection with openssl s_client
|
||||
# b. Parse notAfter from the served certificate
|
||||
# c. Compute days remaining
|
||||
# d. Classify: HEALTHY (silent) | WARNING (≤ CERT_WARN_DAYS)
|
||||
# | CRITICAL (≤ CERT_CRIT_DAYS) | FAILED (no connect / no parse)
|
||||
# 3. Batch by severity — one notification listing all WARNING domains, a separate
|
||||
# one listing all CRITICAL domains
|
||||
#
|
||||
# A domain that fails to connect is reported as FAILED rather than assumed healthy or
|
||||
# assumed expired — an unreachable host and an expiring cert are different problems.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -44,8 +61,10 @@
|
||||
# CERT_TIMEOUT caps each openssl connection attempt. One unreachable domain
|
||||
# does not block the remaining domains.
|
||||
#
|
||||
# Notification Validated
|
||||
# validate_unraid_cmd confirms openssl and notify script are present before use.
|
||||
# openssl Validated
|
||||
# platform_require_cmd confirms openssl is present before any domain is checked — every
|
||||
# check depends on it, so a missing binary is reported as itself rather than as every
|
||||
# domain failing. The notify script is validated separately by the platform adapter.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -83,7 +102,8 @@
|
||||
# Show domain list, warning thresholds, and timeout. Then exit.
|
||||
#
|
||||
# cert_monitor.sh --log
|
||||
# Verbose per-domain output during the run.
|
||||
# Include healthy domains in per-domain output with expiry date and days remaining.
|
||||
# Problems (WARN/CRIT/FAIL) always show with their details regardless of this flag.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
@@ -105,15 +125,11 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
fi
|
||||
|
||||
# Validate openssl — required for all cert checks
|
||||
validate_unraid_cmd \
|
||||
platform_require_cmd \
|
||||
"$(command -v openssl 2>/dev/null || echo /usr/bin/openssl)" \
|
||||
"version" "OpenSSL" \
|
||||
"openssl" || { error "openssl not found — required for certificate checks"; exit 1; }
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
@@ -128,6 +144,8 @@ if [[ ${#CERT_MONITOR_DOMAINS[@]} -eq 0 ]]; then
|
||||
fi
|
||||
|
||||
info "Domains to check: ${#CERT_MONITOR_DOMAINS[@]}"
|
||||
log "$ICON_GEAR Config: warn=${CERT_WARN_DAYS}d crit=${CERT_CRIT_DAYS}d timeout=${CERT_TIMEOUT}s"
|
||||
log "$ICON_GEAR Domains: ${CERT_MONITOR_DOMAINS[*]}"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — results shown but no notifications sent"
|
||||
|
||||
@@ -161,38 +179,23 @@ check_cert() {
|
||||
local domain="$1"
|
||||
local port="${2:-443}"
|
||||
|
||||
local expiry_str
|
||||
expiry_str=$(echo | timeout "$CERT_TIMEOUT" openssl s_client \
|
||||
-connect "${domain}:${port}" \
|
||||
-servername "$domain" \
|
||||
2>/dev/null | openssl x509 -noout -enddate 2>/dev/null | cut -d= -f2)
|
||||
|
||||
if [[ -z "$expiry_str" ]]; then
|
||||
error "$ICON_CERT $domain — could not retrieve certificate (unreachable or no TLS)"
|
||||
if ! check_cert_expiry "$domain" "$port" "$CERT_TIMEOUT"; then
|
||||
if [[ -n "$_CERT_EXPIRY_RAW" ]]; then
|
||||
error "$ICON_CERT $domain — could not parse expiry date: $_CERT_EXPIRY_RAW"
|
||||
else
|
||||
error "$ICON_CERT $domain — could not retrieve certificate (unreachable or no TLS)"
|
||||
fi
|
||||
return 3
|
||||
fi
|
||||
|
||||
local expiry_epoch
|
||||
expiry_epoch=$(date -d "$expiry_str" +%s 2>/dev/null)
|
||||
|
||||
if [[ -z "$expiry_epoch" ]]; then
|
||||
error "$ICON_CERT $domain — could not parse expiry date: $expiry_str"
|
||||
return 3
|
||||
fi
|
||||
|
||||
local now days_remaining expiry_display
|
||||
now=$(date +%s)
|
||||
days_remaining=$(( (expiry_epoch - now) / 86400 ))
|
||||
expiry_display=$(date -d "$expiry_str" '+%Y-%m-%d' 2>/dev/null)
|
||||
|
||||
if [[ "$days_remaining" -le "$CERT_CRIT_DAYS" ]]; then
|
||||
error "$ICON_CERT $domain — CRITICAL: ${days_remaining} days remaining (expires $expiry_display)"
|
||||
if [[ "$_CERT_DAYS" -le "$CERT_CRIT_DAYS" ]]; then
|
||||
error "$ICON_CERT $domain — CRITICAL: ${_CERT_DAYS} days remaining (expires $_CERT_EXPIRY)"
|
||||
return 2
|
||||
elif [[ "$days_remaining" -le "$CERT_WARN_DAYS" ]]; then
|
||||
warn "$ICON_CERT $domain — WARNING: ${days_remaining} days remaining (expires $expiry_display)"
|
||||
elif [[ "$_CERT_DAYS" -le "$CERT_WARN_DAYS" ]]; then
|
||||
warn "$ICON_CERT $domain — WARNING: ${_CERT_DAYS} days remaining (expires $_CERT_EXPIRY)"
|
||||
return 1
|
||||
else
|
||||
log "$ICON_CERT $domain — OK: ${days_remaining} days remaining (expires $expiry_display)"
|
||||
log "$ICON_CERT $domain — OK: ${_CERT_DAYS} days remaining (expires $_CERT_EXPIRY)"
|
||||
return 0
|
||||
fi
|
||||
}
|
||||
@@ -212,12 +215,14 @@ HEALTHY=()
|
||||
WARNING=()
|
||||
CRITICAL=()
|
||||
FAILED=()
|
||||
declare -A DOMAIN_STATUS
|
||||
declare -A DOMAIN_STATUS DOMAIN_DAYS DOMAIN_EXPIRY
|
||||
|
||||
for domain in "${CERT_MONITOR_DOMAINS[@]}"; do
|
||||
[[ -z "$domain" ]] && continue
|
||||
check_cert "$domain"
|
||||
result=$?
|
||||
DOMAIN_DAYS["$domain"]="${_CERT_DAYS:-}"
|
||||
DOMAIN_EXPIRY["$domain"]="${_CERT_EXPIRY:-}"
|
||||
case $result in
|
||||
0) HEALTHY+=("$domain"); DOMAIN_STATUS["$domain"]="OK" ;;
|
||||
1) WARNING+=("$domain"); DOMAIN_STATUS["$domain"]="WARN" ;;
|
||||
@@ -255,13 +260,14 @@ echo " $ICON_SUCCESS Healthy: ${#HEALTHY[@]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${#FAILED[@]} — unreachable"
|
||||
echo ""
|
||||
|
||||
# Per-domain results — only show problems, healthy ones stay in log()
|
||||
# Per-domain results — problems always shown with days remaining; healthy only with --log
|
||||
for domain in "${CERT_MONITOR_DOMAINS[@]}"; do
|
||||
[[ -z "$domain" ]] && continue
|
||||
_days="${DOMAIN_DAYS[$domain]:-?}" _exp="${DOMAIN_EXPIRY[$domain]:-unknown}"
|
||||
case "${DOMAIN_STATUS[$domain]:-UNKN}" in
|
||||
OK) log " $ICON_SUCCESS $domain — healthy" ;;
|
||||
WARN) warn " $ICON_WARN $domain — warning" ;;
|
||||
CRIT) echo " $ICON_ERROR $domain — CRITICAL" ;;
|
||||
OK) log " $ICON_SUCCESS $domain — healthy (${_days}d, expires ${_exp})" ;;
|
||||
WARN) warn " $ICON_WARN $domain — warning (${_days}d, expires ${_exp})" ;;
|
||||
CRIT) echo " $ICON_ERROR $domain — CRITICAL (${_days}d, expires ${_exp})" ;;
|
||||
FAIL) echo " $ICON_ERROR $domain — unreachable" ;;
|
||||
esac
|
||||
done
|
||||
@@ -278,5 +284,24 @@ else
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# ── Write JSON status cache ───────────────────────────────────────────────────
|
||||
_CERT_CACHE_FILE="$STATE_DIR/cert_status.json"
|
||||
{
|
||||
printf '{"checked_at":%d,"host":"%s","warn_days":%d,"crit_days":%d,"dry_run":%s,"domains":[\n' \
|
||||
"$(date +%s)" "$MY_ID" "$CERT_WARN_DAYS" "$CERT_CRIT_DAYS" \
|
||||
"$([[ $DRY_RUN == true ]] && echo true || echo false)"
|
||||
_first=true
|
||||
for _d in "${CERT_MONITOR_DOMAINS[@]}"; do
|
||||
[[ -z "$_d" ]] && continue
|
||||
[[ "$_first" != true ]] && printf ','
|
||||
_first=false
|
||||
_days="${DOMAIN_DAYS[$_d]:-null}"
|
||||
_exp="${DOMAIN_EXPIRY[$_d]:-}"
|
||||
printf '{"domain":"%s","status":"%s","days":%s,"expires":"%s"}\n' \
|
||||
"$_d" "${DOMAIN_STATUS[$_d]:-UNKN}" "$_days" "$_exp"
|
||||
done
|
||||
printf ']}\n'
|
||||
} > "$_CERT_CACHE_FILE" 2>/dev/null
|
||||
|
||||
[[ ${#CRITICAL[@]} -gt 0 || ${#FAILED[@]} -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
Regular → Executable
+58
-27
@@ -18,6 +18,41 @@
|
||||
# configuration issue. Silent on clean runs.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Reachability — Emby API responding; unreachable exits cleanly rather than
|
||||
# reporting an empty library as a real result
|
||||
# 2. Server info and uptime
|
||||
# 3. Active sessions — count, and the transcode-to-direct-play ratio
|
||||
# 4. Library counts — movies, episodes, songs
|
||||
# 5. Activity history over the last EMBY_REPORT_DAYS
|
||||
# 6. Top EMBY_REPORT_TOP_N items and most active users
|
||||
# 7. Ramdisk transcode status, read from the shared transcode state
|
||||
#
|
||||
# Every figure is queried fresh. The only notification is the transcode-ratio warning;
|
||||
# everything else is report output.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# No Persistent State
|
||||
# Every report is generated fresh from the Emby API. No local database, no
|
||||
# incremental tracking. A missed run leaves no gap — the next run simply
|
||||
# covers its own window.
|
||||
#
|
||||
# Silent When Healthy
|
||||
# The report goes to Discord/notification as a summary. Transcode alerts are
|
||||
# the only proactive notification — high transcode ratios may indicate a
|
||||
# misconfigured client that needs attention before it becomes a performance issue.
|
||||
#
|
||||
# Section Independence
|
||||
# Each report section (sessions, library, activity, top content) guards its own
|
||||
# API calls. A failure in one section does not abort the others — the report
|
||||
# produces partial output rather than nothing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -35,9 +70,6 @@
|
||||
# detect_hosts() aliases HOST*_EMBY_URL and HOST*_EMBY_API_KEY → EMBY_URL / EMBY_API_KEY.
|
||||
# Each server reports on its own Emby instance automatically.
|
||||
#
|
||||
# Notification Validated
|
||||
# validate_unraid_cmd confirms the notify script is present before use.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
@@ -111,10 +143,6 @@ if ! command -v jq >/dev/null 2>&1; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
@@ -124,7 +152,7 @@ detect_hosts
|
||||
require_var EMBY_URL
|
||||
require_var EMBY_API_KEY
|
||||
|
||||
log "Emby: $EMBY_URL"
|
||||
log "$ICON_GEAR Config: url=${EMBY_URL} period=${EMBY_REPORT_DAYS}d top=${EMBY_REPORT_TOP_N}"
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — API queried but no notification sent"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -190,7 +218,7 @@ SYSTEM_INFO=$(emby_api "System/Info" 2>/dev/null) || {
|
||||
|
||||
SERVER_NAME=$(echo "$SYSTEM_INFO" | jq -r '.ServerName // "Unknown"' 2>/dev/null)
|
||||
SERVER_VERSION=$(echo "$SYSTEM_INFO" | jq -r '.Version // "Unknown"' 2>/dev/null)
|
||||
log "Connected to: $SERVER_NAME (v$SERVER_VERSION)"
|
||||
echo "$ICON_EMBY Connected to: $SERVER_NAME (v$SERVER_VERSION) ✅"
|
||||
|
||||
# ── Active Sessions ───────────────────────────────────────────────────────────────────────────
|
||||
echo "━━━ $ICON_EMBY Active Sessions ━━━"
|
||||
@@ -203,6 +231,9 @@ TRANSCODE_NOW=$(echo "$SESSIONS" | \
|
||||
2>/dev/null || echo 0)
|
||||
DIRECT_NOW=$(( ACTIVE_COUNT - TRANSCODE_NOW ))
|
||||
|
||||
TRANSCODE_PCT=0
|
||||
[[ "$ACTIVE_COUNT" -gt 0 ]] && TRANSCODE_PCT=$(( TRANSCODE_NOW * 100 / ACTIVE_COUNT ))
|
||||
|
||||
echo " $ICON_EMBY Active streams: $ACTIVE_COUNT"
|
||||
echo " $ICON_EMBY Direct play: $DIRECT_NOW"
|
||||
echo " $ICON_EMBY Transcoding: $TRANSCODE_NOW"
|
||||
@@ -216,6 +247,7 @@ if [[ "$ACTIVE_COUNT" -gt 0 ]]; then
|
||||
" \(.UserName // "Unknown") → \(.NowPlayingItem.Name // "Unknown") [\(if .TranscodingInfo != null then "transcode" else "direct" end)]"
|
||||
' 2>/dev/null || true
|
||||
fi
|
||||
log "$ICON_EMBY Sessions: $ACTIVE_COUNT active ($DIRECT_NOW direct / $TRANSCODE_NOW transcode)"
|
||||
echo ""
|
||||
|
||||
# ── Library Stats ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -229,6 +261,7 @@ SONG_COUNT=$(echo "$ITEMS" | jq '.SongCount // 0' 2>/dev/null || echo 0)
|
||||
echo " $ICON_EMBY Movies: $MOVIE_COUNT"
|
||||
echo " $ICON_EMBY Episodes: $EPISODE_COUNT"
|
||||
echo " $ICON_EMBY Songs: $SONG_COUNT"
|
||||
log "$ICON_EMBY Library: $MOVIE_COUNT movies $EPISODE_COUNT episodes $SONG_COUNT songs"
|
||||
echo ""
|
||||
|
||||
# ── Activity History ──────────────────────────────────────────────────────────────────────────
|
||||
@@ -240,30 +273,28 @@ REPORT_START=$(date -d "${EMBY_REPORT_DAYS} days ago" '+%Y-%m-%dT00:00:00.000Z')
|
||||
ACTIVITY=$(emby_api "System/ActivityLog/Entries?MinDate=${REPORT_START}&Limit=1000" \
|
||||
2>/dev/null) || { warn "Could not fetch activity log"; ACTIVITY="{}"; }
|
||||
|
||||
# Emby 4.9 event types changed: VideoPlayback → playback.stop (completed session)
|
||||
TOTAL_PLAYS=$(echo "$ACTIVITY" | \
|
||||
jq '[.Items // [] | .[] | select(.Type == "VideoPlayback" or .Type == "AudioPlayback")] | length' \
|
||||
2>/dev/null || echo 0)
|
||||
|
||||
TRANSCODE_PLAYS=$(echo "$ACTIVITY" | \
|
||||
jq '[.Items // [] | .[] | select(.Type == "VideoPlaybackUnplugged" or
|
||||
(.Type == "VideoPlayback" and (.Overview // "" | contains("Transcode"))))] | length' \
|
||||
jq '[.Items // [] | .[] | select(.Type == "playback.stop")] | length' \
|
||||
2>/dev/null || echo 0)
|
||||
|
||||
echo " $ICON_EMBY Total play events: $TOTAL_PLAYS"
|
||||
|
||||
if [[ "$TOTAL_PLAYS" -gt 0 ]]; then
|
||||
TRANSCODE_PCT=$(awk "BEGIN {printf \"%.0f\", ($TRANSCODE_PLAYS / $TOTAL_PLAYS) * 100}")
|
||||
DIRECT_PCT=$(( 100 - TRANSCODE_PCT ))
|
||||
echo " $ICON_EMBY Direct play: ~${DIRECT_PCT}%"
|
||||
echo " $ICON_EMBY Transcoded: ~${TRANSCODE_PCT}%"
|
||||
fi
|
||||
[[ "$TOTAL_PLAYS" -gt 0 ]] && log "$ICON_EMBY Activity: $TOTAL_PLAYS plays"
|
||||
echo ""
|
||||
|
||||
# ── Top Content ───────────────────────────────────────────────────────────────────────────────
|
||||
echo "━━━ $ICON_EMBY Top ${EMBY_REPORT_TOP_N} Content ━━━"
|
||||
|
||||
TOP_ITEMS=$(emby_api "Items?SortBy=DatePlayed&SortOrder=Descending&Limit=${EMBY_REPORT_TOP_N}&Recursive=true&Fields=Overview&IncludeItemTypes=Movie,Episode" \
|
||||
2>/dev/null) || { warn "Could not fetch top content"; TOP_ITEMS="{}"; }
|
||||
# Emby 4.9 requires a UserId on the Items endpoint — fetch admin ID first
|
||||
EMBY_ADMIN_ID=$(emby_api "Users" 2>/dev/null | \
|
||||
jq -r '[.[] | select(.Policy.IsAdministrator == true)] | first | .Id // empty' 2>/dev/null)
|
||||
if [[ -z "$EMBY_ADMIN_ID" ]]; then
|
||||
warn "Could not resolve admin UserId — skipping top content"
|
||||
TOP_ITEMS="{}"
|
||||
else
|
||||
TOP_ITEMS=$(emby_api "Users/${EMBY_ADMIN_ID}/Items?SortBy=DatePlayed&SortOrder=Descending&Limit=${EMBY_REPORT_TOP_N}&Recursive=true&Fields=Overview&IncludeItemTypes=Movie,Episode" \
|
||||
2>/dev/null) || { warn "Could not fetch top content"; TOP_ITEMS="{}"; }
|
||||
fi
|
||||
|
||||
TOP_COUNT=$(echo "$TOP_ITEMS" | jq '.Items // [] | length' 2>/dev/null || echo 0)
|
||||
if [[ "$TOP_COUNT" -gt 0 ]]; then
|
||||
@@ -298,7 +329,7 @@ echo ""
|
||||
echo "━━━ $ICON_RAM Transcode Status ━━━"
|
||||
if mountpoint -q "$RAMDISK_PATH" 2>/dev/null; then
|
||||
RAMDISK_USED_KB=$(df "$RAMDISK_PATH" --output=used 2>/dev/null | tail -1 | tr -d ' ')
|
||||
RAMDISK_USED_GB=$(awk "BEGIN {printf \"%.2f\", ${RAMDISK_USED_KB:-0} / 1048576}")
|
||||
RAMDISK_USED_GB=$(kb_to_gb "$RAMDISK_USED_KB")
|
||||
SYMLINK=$(readlink "$TRANSCODE_LINK" 2>/dev/null || echo "unknown")
|
||||
echo " $ICON_RAM Ramdisk usage: ${RAMDISK_USED_GB}GB / ${RAMDISK_SIZE:-8G}"
|
||||
echo " $ICON_LINK Symlink target: $SYMLINK"
|
||||
@@ -318,7 +349,7 @@ END=$(date +%s)
|
||||
echo "━━━━━ $ICON_SUMMARY EMBY REPORT SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_EMBY Server: $SERVER_NAME (v$SERVER_VERSION)"
|
||||
echo "$ICON_EMBY Active: $ACTIVE_COUNT streams ($DIRECT_NOW direct / $TRANSCODE_NOW transcode)"
|
||||
echo "$ICON_EMBY Active: $ACTIVE_COUNT streams ($DIRECT_NOW direct / $TRANSCODE_NOW transcode, ${TRANSCODE_PCT}%)"
|
||||
echo "$ICON_EMBY Library: $MOVIE_COUNT movies $EPISODE_COUNT episodes $SONG_COUNT songs"
|
||||
echo "$ICON_EMBY Period: $TOTAL_PLAYS play events in last ${EMBY_REPORT_DAYS} days"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
@@ -326,7 +357,7 @@ echo "━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
# Only notify on issues — high transcode rate may indicate config problem
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
if [[ "$TOTAL_PLAYS" -gt 10 && "${TRANSCODE_PCT:-0}" -gt 80 ]]; then
|
||||
if [[ "$TOTAL_PLAYS" -gt 10 && "$TRANSCODE_PCT" -gt 80 ]]; then
|
||||
notify "Emby report on $(hostname) — high transcode rate: ${TRANSCODE_PCT}% of $TOTAL_PLAYS plays — check direct play config" \
|
||||
"Emby Report" "warning"
|
||||
fi
|
||||
|
||||
+54
-16
@@ -24,6 +24,43 @@
|
||||
# Safe to share with mesh members — contains no API keys, passwords, or SSH keys.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Read-Only, No Network Calls
|
||||
# All data comes from conf files — no SSH, no API calls, no pings. The output
|
||||
# is always instant and never fails due to a node being unreachable. This makes
|
||||
# it safe to run at any time without side effects.
|
||||
#
|
||||
# Scales Automatically
|
||||
# Iterates all defined HOST* vars rather than a hardcoded list. Adding a new
|
||||
# node to master.conf/host*.conf makes it appear in the output immediately.
|
||||
#
|
||||
# Safe to Share
|
||||
# Output contains only identity and coverage configuration — no API keys,
|
||||
# no passwords, no SSH keys. The report can be shared with other mesh members
|
||||
# without exposing secrets.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# No External Dependencies
|
||||
# Reads only from already-sourced conf files. No curl, no ssh, no docker —
|
||||
# nothing that can fail, hang, or require credentials.
|
||||
#
|
||||
# Empty Mesh Guard
|
||||
# collect_hosts() populates ALL_HOST_IDS — if no HOST* vars are defined the
|
||||
# output sections iterate over an empty array and exit cleanly.
|
||||
#
|
||||
# No Root, No Lock, No detect_hosts — Deliberate
|
||||
# This is the one script in the ecosystem that intentionally omits all three, and
|
||||
# they should not be added. It writes nothing, so there is no state for a lock to
|
||||
# protect and no privileged operation to justify a root gate. It reports on every
|
||||
# node rather than acting as one, so detect_hosts() would narrow it to this host's
|
||||
# aliases — the opposite of what it is for. Every HOST* var is read directly instead.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -71,6 +108,7 @@ format_delay() {
|
||||
}
|
||||
|
||||
collect_hosts
|
||||
log "$ICON_HOST Hosts: ${ALL_HOST_IDS[*]} (${#ALL_HOST_IDS[@]} in mesh)"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
@@ -127,6 +165,7 @@ for h in "${ALL_HOST_IDS[@]}"; do
|
||||
email_var="${h}_OWNER_EMAIL"; email="${!email_var:-(not set)}"
|
||||
printf " %-${W_HOST}s %-${W_SERVER}s %-${W_OWNER}s %s\n" \
|
||||
"$h" "$server" "$owner" "$email"
|
||||
log "$h: server=$server owner=$owner email=$email"
|
||||
done
|
||||
|
||||
# ── Protected Services ────────────────────────────────────────────────────────────────────────
|
||||
@@ -140,30 +179,29 @@ for covered in "${ALL_HOST_IDS[@]}"; do
|
||||
covered_owner_var="${covered}_OWNER"; covered_owner="${!covered_owner_var:-$covered}"
|
||||
covered_email_var="${covered}_OWNER_EMAIL"; covered_email="${!covered_email_var:-}"
|
||||
|
||||
# Find who covers this host and collect their tiers
|
||||
# Each host defines its own recovery profile — read directly from covered host's vars
|
||||
any_tiers=false
|
||||
declare -A tier_containers=()
|
||||
declare -A tier_delays=()
|
||||
covered_by=""
|
||||
|
||||
for covering in "${ALL_HOST_IDS[@]}"; do
|
||||
[[ "$covering" == "$covered" ]] && continue
|
||||
for tier in 1 2 3 4; do
|
||||
containers=$(get_array "FALLBACK_${covered}_TIER${tier}")
|
||||
[[ -z "$containers" ]] && continue
|
||||
any_tiers=true
|
||||
tier_containers[$tier]="$containers"
|
||||
delay_var="${covered}_TIER${tier}_DELAY"
|
||||
tier_delays[$tier]="${!delay_var:-0}"
|
||||
done
|
||||
|
||||
for tier in 1 2 3 4; do
|
||||
containers=$(get_array "FALLBACK_${covering}_COVERS_${covered}_TIER${tier}")
|
||||
[[ -z "$containers" ]] && continue
|
||||
any_tiers=true
|
||||
tier_containers[$tier]="$containers"
|
||||
delay_var="${covered}_TIER${tier}_DELAY"
|
||||
tier_delays[$tier]="${!delay_var:-0}"
|
||||
done
|
||||
|
||||
[[ "$any_tiers" == true ]] && {
|
||||
if [[ "$any_tiers" == true ]]; then
|
||||
for covering in "${ALL_HOST_IDS[@]}"; do
|
||||
[[ "$covering" == "$covered" ]] && continue
|
||||
covering_owner_var="${covering}_OWNER"
|
||||
covering_owner="${!covering_owner_var:-$covering}"
|
||||
covered_by="$covering ($covering_owner)"
|
||||
}
|
||||
done
|
||||
covered_by="${covered_by:+$covered_by, }$covering ($covering_owner)"
|
||||
done
|
||||
fi
|
||||
|
||||
[[ "$any_tiers" == false ]] && continue
|
||||
|
||||
|
||||
Regular → Executable
+46
-30
@@ -19,6 +19,45 @@
|
||||
# from master.conf if dynamix.cfg is not found.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Validate smartctl is present (platform_require_cmd)
|
||||
# 2. Resolve temperature thresholds — dynamix.cfg first, master.conf as fallback
|
||||
# 3. Enumerate drives, skipping anything in SMART_IGNORE_DRIVES
|
||||
# 4. Per drive, read live SMART attributes and evaluate:
|
||||
# overall status FAILED → critical
|
||||
# Reallocated_Sector_Ct > 0 → concerning
|
||||
# Current_Pending_Sector > 0 → concerning
|
||||
# Offline_Uncorrectable > 0 → critical
|
||||
# Temperature_Celsius vs warn/crit thresholds
|
||||
# Power_On_Hours → informational only
|
||||
# NVMe drives expose different attribute names — detected and mapped automatically.
|
||||
# 5. Report; notify only when something crosses a threshold. Silent when all pass.
|
||||
#
|
||||
# Read-only throughout — this queries attributes the drive already maintains and never
|
||||
# starts a self-test. Running one is smart_long_test.sh's job.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Live Queries, No Persistent State
|
||||
# Every run queries smartctl directly — no cached attribute history, no trend
|
||||
# tracking. Each report is a snapshot of current drive health. This keeps the
|
||||
# script simple and the output always current.
|
||||
#
|
||||
# Threshold Parity With unRAID Dashboard
|
||||
# Temperature thresholds are read from dynamix.cfg — the same values unRAID
|
||||
# uses on its own dashboard. A consistent threshold means no conflicting alerts
|
||||
# between this script and the built-in unRAID warnings.
|
||||
#
|
||||
# Silent When Healthy
|
||||
# No output, no notification on a clean run. The absence of a report is the
|
||||
# confirmation that all drives passed. Noise from weekly healthy runs would
|
||||
# erode attention to the reports that matter.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -36,8 +75,9 @@
|
||||
# Reads hot/max/hotssd/maxssd from dynamix.cfg so smart_health.sh and unRAID's
|
||||
# dashboard use the same thresholds. Falls back to master.conf values if not found.
|
||||
#
|
||||
# Notifications Validated
|
||||
# validate_unraid_cmd confirms smartctl and notify script are present before use.
|
||||
# smartctl Validated
|
||||
# platform_require_cmd confirms smartctl is present before any drive is queried. The
|
||||
# notify script is validated separately by the platform adapter.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -94,7 +134,7 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
fi
|
||||
|
||||
# Validate smartctl — required for all drive checks
|
||||
validate_unraid_cmd \
|
||||
platform_require_cmd \
|
||||
"$(command -v smartctl 2>/dev/null || echo /usr/bin/smartctl)" \
|
||||
"--version" "smartmontools" \
|
||||
"smartctl" || {
|
||||
@@ -104,10 +144,6 @@ validate_unraid_cmd \
|
||||
exit 1
|
||||
}
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
@@ -141,11 +177,7 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
for drive in /dev/sd? /dev/nvme?; do
|
||||
[[ ! -e "$drive" ]] && continue
|
||||
drive_name=$(basename "$drive")
|
||||
ignored=false
|
||||
for ignore in "${SMART_IGNORE_DRIVES[@]}"; do
|
||||
[[ "$drive_name" == "$ignore" ]] && ignored=true && break
|
||||
done
|
||||
if [[ "$ignored" == true ]]; then
|
||||
if is_in_list "$drive_name" "${SMART_IGNORE_DRIVES[@]}"; then
|
||||
echo " $ICON_WARN $drive — ignored"
|
||||
else
|
||||
echo " $ICON_SMART $drive — would check"
|
||||
@@ -186,18 +218,7 @@ get_drive_temp() {
|
||||
echo "${temp:-}"
|
||||
}
|
||||
|
||||
# Detect if a drive is SSD/NVMe (rotational=0)
|
||||
is_ssd() {
|
||||
local drive="$1"
|
||||
local dev_name
|
||||
dev_name=$(basename "$drive" | sed 's/nvme[0-9]/nvme0/')
|
||||
local rotational="/sys/block/$(basename "$drive")/queue/rotational"
|
||||
[[ -f "$rotational" ]] && [[ "$(cat "$rotational" 2>/dev/null)" == "0" ]] && return 0
|
||||
# NVMe is always SSD
|
||||
[[ "$drive" == *nvme* ]] && return 0
|
||||
return 1
|
||||
}
|
||||
|
||||
# is_ssd() — provided by common.sh
|
||||
# ==============================================================================================
|
||||
# ━━━ SMART Health Check ━━━
|
||||
# ==============================================================================================
|
||||
@@ -217,12 +238,7 @@ for drive in /dev/sd? /dev/nvme?; do
|
||||
drive_name=$(basename "$drive")
|
||||
|
||||
# Check ignore list
|
||||
ignored=false
|
||||
for ignore in "${SMART_IGNORE_DRIVES[@]}"; do
|
||||
[[ "$drive_name" == "$ignore" ]] && ignored=true && break
|
||||
done
|
||||
|
||||
if [[ "$ignored" == true ]]; then
|
||||
if is_in_list "$drive_name" "${SMART_IGNORE_DRIVES[@]}"; then
|
||||
log "$drive_name — ignored (SMART_IGNORE_DRIVES)"
|
||||
DRIVES_SKIP+=("$drive_name")
|
||||
continue
|
||||
|
||||
@@ -34,6 +34,24 @@
|
||||
# warnings over the week to show trend severity.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Snapshot for Trend, Not Just Alert
|
||||
# Exhaustion events are rarely instant — they build over hours or days.
|
||||
# Logging every 6 hours builds a trend that weekly_health_digest.sh can
|
||||
# surface as a warning count, catching gradual pressure before it becomes
|
||||
# an outage.
|
||||
#
|
||||
# Bounded Log Size
|
||||
# Log entries are trimmed to TUNING_LOG_RETENTION days on every write.
|
||||
# The log never grows unbounded regardless of how long the server runs.
|
||||
#
|
||||
# Silent When Healthy
|
||||
# No output, no notification on a clean run. Threshold breach is the only
|
||||
# signal — routine snapshots below the threshold produce nothing.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -47,9 +65,6 @@
|
||||
# Trim uses tmp file + mv — partial writes during log rotation cannot corrupt
|
||||
# the accumulated history.
|
||||
#
|
||||
# Notification Validated
|
||||
# validate_unraid_cmd confirms the notify script is present before use.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
@@ -72,7 +87,7 @@
|
||||
# Warn if php-fpm active workers exceed this percentage of PHP_MAX_CHILDREN. (default: 80)
|
||||
#
|
||||
# PHP_MAX_CHILDREN
|
||||
# Maximum php-fpm workers — set by php_fpm_max_children.sh in unRAID_Essentials.
|
||||
# Maximum php-fpm workers — set by php_fpm_max_children.sh in System_Essentials.
|
||||
#
|
||||
# TUNING_MONITOR_LOG
|
||||
# Log file path. (default: DATA_DIR/tuning_monitor.db)
|
||||
@@ -113,16 +128,14 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
# detect_hosts() sets MY_ID — used in warning output
|
||||
detect_hosts
|
||||
|
||||
log "$ICON_GEAR Config: inotify-warn=${INOTIFY_WARN_PCT:-80}% php-fpm-warn=${PHP_FPM_WARN_PCT:-80}% max-workers=${PHP_MAX_CHILDREN:-250} retention=${TUNING_LOG_RETENTION:-30}d"
|
||||
|
||||
DATE=$(date '+%Y-%m-%d')
|
||||
TIME=$(date '+%H:%M')
|
||||
INOTIFY_WARN=0
|
||||
@@ -196,7 +209,7 @@ fi
|
||||
# ==============================================================================================
|
||||
PHPFPM_MAX="${PHP_MAX_CHILDREN:-250}"
|
||||
|
||||
PHPFPM_ACTIVE=$(ps aux 2>/dev/null | grep -c "php-fpm: pool" || echo 0)
|
||||
PHPFPM_ACTIVE=$(ps aux 2>/dev/null | grep -c "php-fpm: pool" || true)
|
||||
PHPFPM_ACTIVE="${PHPFPM_ACTIVE//[^0-9]/}"
|
||||
PHPFPM_ACTIVE="${PHPFPM_ACTIVE:-0}"
|
||||
|
||||
@@ -245,4 +258,5 @@ fi
|
||||
echo "${DATE}|${TIME}|${INOTIFY_USED}|${INOTIFY_LIMIT}|${INOTIFY_PCT}|${INOTIFY_WARN}|${PHPFPM_ACTIVE}|${PHPFPM_MAX}|${PHPFPM_PCT}|${PHPFPM_WARN}" \
|
||||
>> "$TUNING_MONITOR_LOG"
|
||||
|
||||
echo "Snapshot written: inotify ${INOTIFY_PCT}% php-fpm ${PHPFPM_PCT}%"
|
||||
echo "Snapshot written: inotify ${INOTIFY_PCT}% php-fpm ${PHPFPM_PCT}%"
|
||||
log "Entry: ${DATE}|${TIME}|${INOTIFY_USED}/${INOTIFY_LIMIT}(${INOTIFY_PCT}%,warn=${INOTIFY_WARN})|${PHPFPM_ACTIVE}/${PHPFPM_MAX}(${PHPFPM_PCT}%,warn=${PHPFPM_WARN})"
|
||||
|
||||
Executable
+93
@@ -0,0 +1,93 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================== Uptime Report =============================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# The weekly read of what Tools/uptime_probe.sh has been recording every minute: anything down
|
||||
# right now, and anything that was not perfect over the last seven days. Runs in the Sunday
|
||||
# Morning Coffee Report.
|
||||
#
|
||||
# Silent on a clean week. A report that always says something is a report nobody reads, so this
|
||||
# prints nothing and notifies nothing when every domain was 100%.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# uptime_probe.php --report exits 1 when it has something to say and 0 when it does not, so the
|
||||
# decision to notify is the exit code rather than this script parsing the text it just printed.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Silence is the normal output.
|
||||
# A report that always says something is a report nobody reads. A perfect week prints nothing
|
||||
# and notifies nothing, so anything that does appear in the Sunday report is worth the glance.
|
||||
#
|
||||
# The exit code is the decision, not the text.
|
||||
# uptime_probe.php --report exits 1 when it has something to say and 0 when it does not. This
|
||||
# script never parses the output it just printed to work out whether to notify — a report whose
|
||||
# wording changed would otherwise silently stop notifying.
|
||||
#
|
||||
# It reads; it never probes.
|
||||
# The measurements are already taken, once a minute, by Tools/uptime_probe.sh. Re-probing at
|
||||
# report time would describe Sunday morning rather than the week being reported on.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Read-only. Reads the stored history and prints; records nothing, and cannot alter the data it
|
||||
# is reporting on.
|
||||
#
|
||||
# UPTIME_PROBE_ENABLED gates the whole run — with the probe off there is no history worth
|
||||
# reporting, and this says nothing rather than reporting an empty week as a perfect one.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# uptime_report.sh the weekly read. Silent when every domain was 100%.
|
||||
#
|
||||
# Called from COFFEE_REPORT_SCRIPTS; takes no arguments and has no other mode. For live figures
|
||||
# or a per-domain table, use Tools/uptime_probe.sh --status.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# UPTIME_PROBE_ENABLED nothing here runs when the probe is switched off
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
acquire_lock
|
||||
detect_hosts
|
||||
|
||||
if [[ "${UPTIME_PROBE_ENABLED:-true}" == "false" ]]; then
|
||||
log "$ICON_GEAR Uptime probe disabled — nothing to report"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
REPORT="$(php "$SCRIPT_DIR/../Plugin/unraid/Tools/uptime_probe.php" --report 2>/dev/null)"
|
||||
RC=$?
|
||||
|
||||
if [[ $RC -eq 0 || -z "$REPORT" ]]; then
|
||||
echo "$ICON_DONE All monitored domains at 100% this week ✅"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "$REPORT"
|
||||
|
||||
DOWN_COUNT=$(grep -c "DOWN" <<< "$REPORT" || true)
|
||||
if [[ "$DOWN_COUNT" -gt 0 ]]; then
|
||||
notify "$DOWN_COUNT domain(s) currently unreachable on $(hostname)" "Uptime" "warning"
|
||||
else
|
||||
notify "Some domains had downtime this week on $(hostname)" "Uptime" "normal"
|
||||
fi
|
||||
exit 0
|
||||
@@ -31,7 +31,7 @@
|
||||
#
|
||||
# Data sources (reads only):
|
||||
# FALLBACK_STATE_FILE — current fallback state
|
||||
# SYS_WATCHDOG_FAILED_FILE — container skip list (manual intervention needed)
|
||||
# DOCKER_WATCHDOG_FAILED_FILE — container skip list (manual intervention needed)
|
||||
# WATCHDOG_STATE_FILE — active container watchdog strikes
|
||||
# SYS_WATCHDOG_STATE_FILE — active system watchdog strikes
|
||||
# BANDWIDTH_LOG — yesterday's transfer totals
|
||||
@@ -40,6 +40,24 @@
|
||||
# RAMDISK_PATH / TRANSCODE_LINK — current transcode location and usage
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Aggregator, Not Generator
|
||||
# This script reads state files that other scripts maintain. It never produces
|
||||
# health data itself — it only presents what is already there. Each source
|
||||
# script remains responsible for its own state; this script is the envelope.
|
||||
#
|
||||
# Profile-Driven Notification
|
||||
# The cron schedule never changes. The DIGEST_PROFILE in master.conf controls
|
||||
# when notifications actually send — switching from daily noise to weekly
|
||||
# summaries is a one-line conf change, not a cron edit.
|
||||
#
|
||||
# Read-Only, No Side Effects
|
||||
# Writes nothing, changes nothing, triggers nothing. Safe to run at any time
|
||||
# for a health snapshot without affecting any running service or state file.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -58,8 +76,10 @@
|
||||
# smart profile produces no output and no notification when nothing worth
|
||||
# reporting is found.
|
||||
#
|
||||
# Notifications Validated
|
||||
# validate_unraid_cmd confirms notify and openssl are present before use.
|
||||
# openssl Validated — Non-Fatal
|
||||
# platform_require_cmd checks openssl and, unlike the other monitors, only warns if it
|
||||
# is missing: the SSL section is skipped and the rest of the digest still runs. The
|
||||
# notify script is validated separately by the platform adapter.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
@@ -126,12 +146,8 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
validate_unraid_cmd \
|
||||
platform_require_cmd \
|
||||
"$(command -v openssl 2>/dev/null || echo /usr/bin/openssl)" \
|
||||
"version" "OpenSSL" \
|
||||
"openssl" || warn "openssl not found — SSL cert checks will be skipped"
|
||||
@@ -141,6 +157,9 @@ acquire_lock
|
||||
# detect_hosts() sets MY_ID and aliases all host-specific vars used in this report
|
||||
detect_hosts
|
||||
|
||||
log "$ICON_GEAR Config: profile=${DIGEST_PROFILE} day=${DIGEST_DAY}"
|
||||
log "$ICON_GEAR Smart triggers: watchdog=${DIGEST_SMART_ON_WATCHDOG} fallback=${DIGEST_SMART_ON_FALLBACK} cert=${DIGEST_SMART_ON_CERT_WARN} bandwidth=${DIGEST_SMART_ON_BANDWIDTH}"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — report generated but no notification sent"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -176,7 +195,7 @@ case "$DIGEST_PROFILE" in
|
||||
SHOULD_SEND=true
|
||||
log "Profile: weekly — today is $DIGEST_DAY — will send"
|
||||
else
|
||||
log "Profile: weekly — today is $TODAY_NAME, digest day is $DIGEST_DAY — skipping"
|
||||
echo "Profile: weekly — today is $TODAY_NAME, digest day is $DIGEST_DAY — no-op"
|
||||
exit 0
|
||||
fi
|
||||
;;
|
||||
@@ -198,7 +217,8 @@ FINDINGS=() # notable but not critical
|
||||
ISSUES=() # need attention
|
||||
DIGEST_LINES=() # full report lines
|
||||
|
||||
# ── Failover State ────────────────────────────────────────────────────────────────────────────
|
||||
# ── fallback State ────────────────────────────────────────────────────────────────────────────
|
||||
log "Reading: fallback=$FALLBACK_STATE_FILE skip=$DOCKER_WATCHDOG_FAILED_FILE watchdog=$WATCHDOG_STATE_FILE sys=$SYS_WATCHDOG_STATE_FILE"
|
||||
if [[ -f "$FALLBACK_STATE_FILE" ]]; then
|
||||
FALLBACK_STATE=$(grep "^state=" "$FALLBACK_STATE_FILE" 2>/dev/null | cut -d= -f2)
|
||||
if [[ -n "$FALLBACK_STATE" ]]; then
|
||||
@@ -213,9 +233,9 @@ else
|
||||
fi
|
||||
|
||||
# ── Container Skip List ───────────────────────────────────────────────────────────────────────
|
||||
if [[ -f "$SYS_WATCHDOG_FAILED_FILE" ]] && [[ -s "$SYS_WATCHDOG_FAILED_FILE" ]]; then
|
||||
SKIP_COUNT=$(wc -l < "$SYS_WATCHDOG_FAILED_FILE")
|
||||
SKIP_LIST=$(cat "$SYS_WATCHDOG_FAILED_FILE" | tr '\n' ' ')
|
||||
if [[ -f "$DOCKER_WATCHDOG_FAILED_FILE" ]] && [[ -s "$DOCKER_WATCHDOG_FAILED_FILE" ]]; then
|
||||
SKIP_COUNT=$(wc -l < "$DOCKER_WATCHDOG_FAILED_FILE")
|
||||
SKIP_LIST=$(cat "$DOCKER_WATCHDOG_FAILED_FILE" | tr '\n' ' ')
|
||||
DIGEST_LINES+=("$ICON_NOT_RUNNING Skip list: $SKIP_COUNT containers — $SKIP_LIST")
|
||||
ISSUES+=("Containers on skip list (manual intervention needed): $SKIP_LIST")
|
||||
SHOULD_SEND=true
|
||||
@@ -238,9 +258,9 @@ fi
|
||||
|
||||
# ── System Watchdog Strikes ───────────────────────────────────────────────────────────────────
|
||||
if [[ -f "$SYS_WATCHDOG_STATE_FILE" ]]; then
|
||||
SYS_STRIKES=$(grep -v ":0$" "$SYS_WATCHDOG_STATE_FILE" 2>/dev/null | grep -c ".")
|
||||
SYS_STRIKES=$(grep "^[^=]*:" "$SYS_WATCHDOG_STATE_FILE" 2>/dev/null | grep -cv -E "(:0$|:false$)")
|
||||
if [[ "$SYS_STRIKES" -gt 0 ]]; then
|
||||
SYS_STRIKE_LIST=$(grep -v ":0$" "$SYS_WATCHDOG_STATE_FILE" 2>/dev/null | tr '\n' ' ')
|
||||
SYS_STRIKE_LIST=$(grep "^[^=]*:" "$SYS_WATCHDOG_STATE_FILE" 2>/dev/null | grep -v -E "(:0$|:false$)" | tr '\n' ' ')
|
||||
DIGEST_LINES+=("$ICON_REBOOT_SMART System strikes: $SYS_STRIKES active — $SYS_STRIKE_LIST")
|
||||
FINDINGS+=("System watchdog: $SYS_STRIKES active strikes")
|
||||
[[ "$DIGEST_SMART_ON_WATCHDOG" == true ]] && SHOULD_SEND=true
|
||||
@@ -252,7 +272,7 @@ fi
|
||||
# ── Transcode Ramdisk ─────────────────────────────────────────────────────────────────────────
|
||||
if mountpoint -q "$RAMDISK_PATH" 2>/dev/null; then
|
||||
RAMDISK_USED_KB=$(df "$RAMDISK_PATH" --output=used 2>/dev/null | tail -1 | tr -d ' ')
|
||||
RAMDISK_USED_GB=$(awk "BEGIN {printf \"%.2f\", ${RAMDISK_USED_KB:-0} / 1048576}")
|
||||
RAMDISK_USED_GB=$(kb_to_gb "$RAMDISK_USED_KB")
|
||||
SYMLINK_TARGET=$(readlink "$TRANSCODE_LINK" 2>/dev/null || echo "unknown")
|
||||
DIGEST_LINES+=("$ICON_RAM Transcodes: ${RAMDISK_USED_GB}GB used → $SYMLINK_TARGET")
|
||||
|
||||
@@ -291,7 +311,7 @@ if [[ -f "${BANDWIDTH_LOG:-}" ]] && [[ -s "$BANDWIDTH_LOG" ]]; then
|
||||
YESTERDAY=$(date -d "yesterday" '+%Y-%m-%d')
|
||||
YESTERDAY_BYTES=$(awk -F'|' -v d="$YESTERDAY" '$1==d{sum+=$6} END{print sum+0}' \
|
||||
"$BANDWIDTH_LOG")
|
||||
YESTERDAY_GB=$(awk "BEGIN {printf \"%.2f\", ${YESTERDAY_BYTES:-0} / 1073741824}")
|
||||
YESTERDAY_GB=$(bytes_to_gb "$YESTERDAY_BYTES")
|
||||
YESTERDAY_LARGE=$(awk -F'|' -v d="$YESTERDAY" '$1==d && $7=="LARGE"' \
|
||||
"$BANDWIDTH_LOG" | wc -l)
|
||||
|
||||
@@ -311,18 +331,16 @@ if [[ ${#CERT_MONITOR_DOMAINS[@]} -gt 0 ]] && command -v openssl >/dev/null 2>&1
|
||||
CERT_ISSUES=()
|
||||
for domain in "${CERT_MONITOR_DOMAINS[@]}"; do
|
||||
[[ -z "$domain" ]] && continue
|
||||
expiry_str=$(echo | timeout "${CERT_TIMEOUT:-10}" openssl s_client \
|
||||
-connect "${domain}:443" -servername "$domain" \
|
||||
2>/dev/null | openssl x509 -noout -enddate 2>/dev/null | cut -d= -f2)
|
||||
if [[ -n "$expiry_str" ]]; then
|
||||
expiry_epoch=$(date -d "$expiry_str" +%s 2>/dev/null)
|
||||
days_remaining=$(( (expiry_epoch - $(date +%s)) / 86400 ))
|
||||
if check_cert_expiry "$domain" 443 "${CERT_TIMEOUT:-10}"; then
|
||||
days_remaining="$_CERT_DAYS"
|
||||
if [[ "$days_remaining" -le "${CERT_CRIT_DAYS:-7}" ]]; then
|
||||
CERT_ISSUES+=("$domain: ${days_remaining}d CRITICAL")
|
||||
SHOULD_SEND=true
|
||||
elif [[ "$days_remaining" -le "${CERT_WARN_DAYS:-30}" ]]; then
|
||||
CERT_ISSUES+=("$domain: ${days_remaining}d warning")
|
||||
[[ "$DIGEST_SMART_ON_CERT_WARN" == true ]] && SHOULD_SEND=true
|
||||
else
|
||||
log "$ICON_CERT $domain: ${days_remaining}d remaining ✅"
|
||||
fi
|
||||
fi
|
||||
done
|
||||
@@ -376,4 +394,4 @@ if [[ "$DRY_RUN" == true ]]; then
|
||||
elif [[ "$SHOULD_SEND" == true ]]; then
|
||||
notify "$NOTIFY_MSG" "Health Digest" "$NOTIFY_SEV"
|
||||
echo "Digest sent"
|
||||
fi
|
||||
fi
|
||||
|
||||
@@ -10,8 +10,9 @@
|
||||
# only — system_watchdog.sh handles threshold-based intervention.
|
||||
#
|
||||
# Combines ZFS pool status, ARC statistics, Docker memory usage, and kernel
|
||||
# memory pressure into a single snapshot. Output goes to both console (for User
|
||||
# Scripts output log) and ZFS_REPORT_LOG for week-over-week comparison.
|
||||
# memory pressure into a single snapshot. In normal mode output goes to both
|
||||
# console (for User Scripts output log) and ZFS_REPORT_LOG for week-over-week
|
||||
# comparison. In --dry-run mode, console only — nothing written to the log.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
@@ -23,15 +24,42 @@
|
||||
# ZFS_REPORT_IGNORE_POOLS excluded from the report
|
||||
# (still fully monitored by unRAID — report-only exclusion).
|
||||
# ARC statistics — current ARC vs max, metadata pressure, hit rate.
|
||||
# Warns if ARC utilisation exceeds ZFS_REPORT_ARC_WARN_PCT.
|
||||
# Memory status — total, free, available RAM.
|
||||
# Warns if free < ZFS_REPORT_FREE_WARN_GB or
|
||||
# available < ZFS_REPORT_AVAIL_WARN_GB.
|
||||
# Warns if ARC utilisation exceeds ZFS_REPORT_ARC_WARN_PCT, or if
|
||||
# ARC headroom (max - current) drops below ZFS_REPORT_ARC_FREE_WARN_GB.
|
||||
# Memory status — total, free, available RAM (informational only — see note below).
|
||||
# Warns if available < ZFS_REPORT_AVAIL_WARN_GB.
|
||||
# Docker memory — top ZFS_REPORT_DOCKER_TOP containers by memory usage.
|
||||
# Useful for spotting containers approaching watchdog limits.
|
||||
# Kernel pressure — vmstat snapshot (3 samples).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Informational, Not Interventional
|
||||
# This script reports — it does not act. system_watchdog.sh handles
|
||||
# threshold-based intervention. Keeping the roles separate means the report
|
||||
# is never suppressed by the same logic that triggers remediation.
|
||||
#
|
||||
# Week-Over-Week Comparison
|
||||
# Output is written to ZFS_REPORT_LOG so the same snapshot can be reviewed
|
||||
# across weeks. Memory pressure and ARC creep are slow — a single run is
|
||||
# rarely conclusive; the trend across weeks is what matters.
|
||||
#
|
||||
# Section Independence
|
||||
# Each of the five report sections guards its own data source. ZFS not
|
||||
# available, Docker not responding — those sections skip, the rest still run.
|
||||
# A partial report is more useful than no report.
|
||||
#
|
||||
# ARC Headroom, Not System Free RAM
|
||||
# ZFS ARC deliberately grows to use most of the RAM the system isn't otherwise
|
||||
# using — that's the point of a page cache. System-wide "free" RAM being low is
|
||||
# therefore normal and not a signal of anything, so the memory warning is based
|
||||
# on ARC headroom (ARC_MAX - ARC_CURRENT) instead — how much room ARC itself has
|
||||
# left before it hits its configured ceiling. "Available" RAM (which accounts for
|
||||
# reclaimable cache) is still checked separately as a true system-pressure signal.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
@@ -51,9 +79,6 @@
|
||||
# Docker Stats Timeout
|
||||
# DOCKER_TIMEOUT caps docker stats calls. A hung daemon does not block the report.
|
||||
#
|
||||
# Notification Validated
|
||||
# validate_unraid_cmd confirms the notify script is present before use.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# STATE FILES
|
||||
# ==============================================================================================
|
||||
@@ -82,8 +107,8 @@
|
||||
# ZFS_REPORT_ARC_WARN_PCT
|
||||
# Warn if ARC is using more than this percentage of its configured max. (default: 90)
|
||||
#
|
||||
# ZFS_REPORT_FREE_WARN_GB
|
||||
# Warn if free RAM is below this threshold in GB. (default: 10)
|
||||
# ZFS_REPORT_ARC_FREE_WARN_GB
|
||||
# Warn if ARC headroom (ARC_MAX - ARC_CURRENT) drops below this many GB. (default: 10)
|
||||
#
|
||||
# ZFS_REPORT_AVAIL_WARN_GB
|
||||
# Warn if available RAM is below this threshold in GB. (default: 20)
|
||||
@@ -125,10 +150,6 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
@@ -143,6 +164,7 @@ done
|
||||
|
||||
log "Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
log "Ignoring pools: ${ZFS_REPORT_IGNORE_POOLS[*]:-none}"
|
||||
log "$ICON_GEAR Config: arc-warn=${ZFS_REPORT_ARC_WARN_PCT}% arc-free-warn=${ZFS_REPORT_ARC_FREE_WARN_GB}GB avail-warn=${ZFS_REPORT_AVAIL_WARN_GB}GB docker-top=${ZFS_REPORT_DOCKER_TOP}"
|
||||
|
||||
# Tee output to log file unless dry run
|
||||
if [[ "$DRY_RUN" == false ]]; then
|
||||
@@ -161,7 +183,7 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_ZFS Log file: $ZFS_REPORT_LOG"
|
||||
echo "$ICON_ZFS ARC warn: ${ZFS_REPORT_ARC_WARN_PCT}%"
|
||||
echo "$ICON_MEM Free RAM warn: ${ZFS_REPORT_FREE_WARN_GB}GB"
|
||||
echo "$ICON_ZFS ARC free warn: ${ZFS_REPORT_ARC_FREE_WARN_GB}GB"
|
||||
echo "$ICON_MEM Avail warn: ${ZFS_REPORT_AVAIL_WARN_GB}GB"
|
||||
echo "$ICON_CONTAINERS Docker top: $ZFS_REPORT_DOCKER_TOP"
|
||||
echo "$ICON_ZFS Ignore pools: ${ZFS_REPORT_IGNORE_POOLS[*]:-none}"
|
||||
@@ -239,21 +261,32 @@ echo "━━━ $ICON_ZFS ARC Statistics ━━━"
|
||||
if [[ ! -f /proc/spl/kstat/zfs/arcstats ]]; then
|
||||
warn "ZFS arcstats not available — skipping ARC section"
|
||||
else
|
||||
ARC_MAX=$(cat /sys/module/zfs/parameters/zfs_arc_max 2>/dev/null || \
|
||||
awk '/^c_max / {print $3}' /proc/spl/kstat/zfs/arcstats)
|
||||
# zfs_arc_max reads 0 when it has been left at the default, which is a value rather than a
|
||||
# failure — so the || fallback never fires for the case that actually needs it, exactly like a
|
||||
# grep -c that prints 0 and exits 1. Zero here would reach the ARC_PCT division below, and awk
|
||||
# treats division by zero as fatal: it prints nothing, ARC_PCT comes back empty, and the whole
|
||||
# ARC section reports blanks. c_max is the cap the kernel is really enforcing either way.
|
||||
ARC_MAX=$(cat /sys/module/zfs/parameters/zfs_arc_max 2>/dev/null || echo 0)
|
||||
ARC_MAX="${ARC_MAX//[^0-9]/}"
|
||||
if [[ "${ARC_MAX:-0}" -eq 0 ]]; then
|
||||
ARC_MAX=$(awk '/^c_max / {print $3}' /proc/spl/kstat/zfs/arcstats 2>/dev/null)
|
||||
ARC_MAX="${ARC_MAX:-0}"
|
||||
fi
|
||||
ARC_SIZE=$(awk '/^size / {print $3}' /proc/spl/kstat/zfs/arcstats)
|
||||
ARC_META_USED=$(awk '/^arc_meta_used / {print $3}' /proc/spl/kstat/zfs/arcstats)
|
||||
|
||||
ARC_MAX_GB=$(awk "BEGIN {printf \"%.1f\", $ARC_MAX / 1073741824}")
|
||||
ARC_CUR_GB=$(awk "BEGIN {printf \"%.1f\", $ARC_SIZE / 1073741824}")
|
||||
ARC_META_GB=$(awk "BEGIN {printf \"%.1f\", $ARC_META_USED / 1073741824}")
|
||||
ARC_MAX_GB=$(bytes_to_gb "$ARC_MAX" 1)
|
||||
ARC_CUR_GB=$(bytes_to_gb "$ARC_SIZE" 1)
|
||||
ARC_META_GB=$(bytes_to_gb "$ARC_META_USED" 1)
|
||||
ARC_PCT=$(awk "BEGIN {printf \"%.1f\", $ARC_SIZE * 100 / $ARC_MAX}")
|
||||
ARC_PCT_INT=$(printf "%.0f" "$ARC_PCT")
|
||||
ARC_FREE_GB=$(awk "BEGIN {printf \"%d\", ($ARC_MAX - $ARC_SIZE) / 1073741824}")
|
||||
|
||||
echo " $ICON_ZFS ARC Max: ${ARC_MAX_GB}GB"
|
||||
echo " $ICON_ZFS ARC Current: ${ARC_CUR_GB}GB"
|
||||
echo " $ICON_ZFS ARC Meta Used: ${ARC_META_GB}GB"
|
||||
echo " $ICON_ZFS ARC Utilization: ${ARC_PCT}%"
|
||||
echo " $ICON_ZFS ARC Free: ${ARC_FREE_GB}GB"
|
||||
|
||||
if [[ "$ARC_PCT_INT" -ge "$ZFS_REPORT_ARC_WARN_PCT" ]]; then
|
||||
warn "ARC utilization ${ARC_PCT}% — above ${ZFS_REPORT_ARC_WARN_PCT}% threshold"
|
||||
@@ -262,6 +295,13 @@ else
|
||||
log "ARC utilization ${ARC_PCT}% — within threshold ✅"
|
||||
fi
|
||||
|
||||
if [[ "$ARC_FREE_GB" -lt "$ZFS_REPORT_ARC_FREE_WARN_GB" ]]; then
|
||||
warn "ARC free ${ARC_FREE_GB}GB — below ${ZFS_REPORT_ARC_FREE_WARN_GB}GB threshold"
|
||||
WARNINGS+=("ARC headroom low: ${ARC_FREE_GB}GB")
|
||||
else
|
||||
log "ARC free ${ARC_FREE_GB}GB — within threshold ✅"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
META_MRU_GHOST=$(awk '/^mru_ghost_metadata / {print $3}' \
|
||||
/proc/spl/kstat/zfs/arcstats 2>/dev/null || echo 0)
|
||||
@@ -270,8 +310,8 @@ else
|
||||
META_MISSES=$(awk '/^demand_metadata_misses / {print $3}' \
|
||||
/proc/spl/kstat/zfs/arcstats 2>/dev/null || echo 0)
|
||||
|
||||
MRU_GB=$(awk "BEGIN {printf \"%.2f\", $META_MRU_GHOST / 1073741824}")
|
||||
MFU_GB=$(awk "BEGIN {printf \"%.2f\", $META_MFU_GHOST / 1073741824}")
|
||||
MRU_GB=$(bytes_to_gb "$META_MRU_GHOST")
|
||||
MFU_GB=$(bytes_to_gb "$META_MFU_GHOST")
|
||||
|
||||
echo " $ICON_ZFS MRU Ghost: ${MRU_GB}GB"
|
||||
echo " $ICON_ZFS MFU Ghost: ${MFU_GB}GB"
|
||||
@@ -285,20 +325,12 @@ echo "━━━ $ICON_MEM Memory Status ━━━"
|
||||
FREE_HUMAN=$(free -h | awk '/Mem:/ {print $4}')
|
||||
AVAIL_HUMAN=$(free -h | awk '/Mem:/ {print $7}')
|
||||
TOTAL_HUMAN=$(free -h | awk '/Mem:/ {print $2}')
|
||||
FREE_GB=$(free -g | awk '/Mem:/ {print $4}')
|
||||
AVAIL_GB=$(free -g | awk '/Mem:/ {print $7}')
|
||||
|
||||
echo " $ICON_MEM Total RAM: $TOTAL_HUMAN"
|
||||
echo " $ICON_MEM Free RAM: $FREE_HUMAN"
|
||||
echo " $ICON_MEM Free RAM: $FREE_HUMAN (informational — ARC intentionally uses most of this)"
|
||||
echo " $ICON_MEM Available RAM: $AVAIL_HUMAN"
|
||||
|
||||
if [[ "$FREE_GB" -lt "$ZFS_REPORT_FREE_WARN_GB" ]]; then
|
||||
warn "Free RAM ${FREE_HUMAN} — below ${ZFS_REPORT_FREE_WARN_GB}GB threshold"
|
||||
WARNINGS+=("Low free RAM: ${FREE_HUMAN}")
|
||||
else
|
||||
log "Free RAM ${FREE_HUMAN} — within threshold ✅"
|
||||
fi
|
||||
|
||||
if [[ "$AVAIL_GB" -lt "$ZFS_REPORT_AVAIL_WARN_GB" ]]; then
|
||||
warn "Available RAM ${AVAIL_HUMAN} — below ${ZFS_REPORT_AVAIL_WARN_GB}GB threshold"
|
||||
WARNINGS+=("Low available RAM: ${AVAIL_HUMAN}")
|
||||
@@ -358,4 +390,4 @@ fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
[[ ${#WARNINGS[@]} -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
exit 0
|
||||
|
||||
@@ -0,0 +1,559 @@
|
||||
# Varaverk AI Integration — Design Notes
|
||||
|
||||
**Status: retrieval is built; integration is not.** As of 2026-08-02 the `AI_*` and
|
||||
`HOST*_OLLAMA_*` variables exist in both confs and both templates, and `AI/` holds a working
|
||||
index and query path — see the RAG section at the end of this document and `AI/README-AI.md`.
|
||||
|
||||
Everything else below remains design only. **No Varaverk script consults AI.** Every
|
||||
`AI_ASSIST_*` toggle is false, `AI_CONF_WRITE_ENABLED` is false with an empty whitelist, and
|
||||
host resolution across the mesh is specified but not implemented. Originally captured
|
||||
2026-08-01 so the reasoning survives.
|
||||
|
||||
Ollama itself *is* installed, tuned and verified on HOST1 — `qwen2.5-coder:14b` for
|
||||
generation, `nomic-embed-text` for embeddings, 16k context, pinned to the RTX 3080. See
|
||||
Hardware Budget for measured numbers. That is the substrate, not the integration.
|
||||
|
||||
**Origin:** the RTX 3080 was freed when the Windows gaming VM was retired. It is bound to the
|
||||
`nvidia` driver, not `vfio` — not reserved for passthrough, so there is no VM contention to
|
||||
design around. Two goals at once: somewhere to learn local LLMs, and something Varaverk can
|
||||
genuinely use.
|
||||
|
||||
**Build order** — deliberately lowest-risk first. Each stage must be boring before the next
|
||||
one starts:
|
||||
|
||||
1. Chat assistant / settings helper / onboarding assistant — a wrong answer costs nothing
|
||||
2. Watchdog and discovery context — a wrong answer costs a bad suggestion, still gated
|
||||
3. Cleanup and sync decision aid — closest to destructive, last to be trusted
|
||||
|
||||
The regular script is always the backup, at every stage.
|
||||
|
||||
---
|
||||
|
||||
## The governing principle
|
||||
|
||||
**Varaverk works exactly as well with AI off as with it on.**
|
||||
|
||||
Every script — including scripts written after this lands — is designed and hardened
|
||||
without AI first. AI is added afterwards as enhancement, never as a dependency. A script
|
||||
that cannot do its job when `AI_ENABLED=false` is a broken script, not an AI feature.
|
||||
|
||||
This is the constraint everything else in this document answers to. If a design decision
|
||||
makes AI load-bearing, that decision is wrong.
|
||||
|
||||
Corollary: AI never makes a destructive decision. The session that produced the current
|
||||
safeguard layer (depth guards, strike thresholds, verification-after-write) exists because
|
||||
config values and scan results feed `rm -rf`, `chown -R`, and rsync `--delete`. AI advises
|
||||
at the points where a script currently stops and defers to a human. The deterministic guard
|
||||
still pulls the trigger.
|
||||
|
||||
---
|
||||
|
||||
## Two independent off-switches
|
||||
|
||||
| switch | meaning | source |
|
||||
|---|---|---|
|
||||
| `AI_ENABLED` | intent — do we want AI at all | `master.conf` |
|
||||
| resolver result | availability — is there a reachable host | runtime probe |
|
||||
|
||||
**Both must produce the identical code path when off.** A caller that gets "no AI" from
|
||||
either must run its normal, non-AI logic — not a degraded variant, not a skipped step.
|
||||
|
||||
`AI_ENABLED` follows the fail-closed idiom standardised across the ecosystem:
|
||||
|
||||
```bash
|
||||
[[ "${AI_ENABLED:-false}" != "true" ]] && <normal path>
|
||||
```
|
||||
|
||||
Not `== false`. Anything that isn't exactly `true` means off, so a typo can never switch AI
|
||||
on. (This is the same bug that was fixed in `fallback.sh` — see its `FALLBACK_ENABLED` gate.)
|
||||
|
||||
---
|
||||
|
||||
## Configuration schema
|
||||
|
||||
Follows the existing rule: **thresholds and toggles → `master.conf`; hardware, paths,
|
||||
container names and per-host identity → `host*.conf`.**
|
||||
|
||||
### `host*.conf` — per-host, because only one node actually has the GPU
|
||||
|
||||
```bash
|
||||
# ━━━ Ollama / AI ━━━
|
||||
HOST1_OLLAMA_URL="http://localhost:11434" # empty on nodes without a local Ollama
|
||||
HOST1_OLLAMA_CONTAINER="Ollama" # for docker_watchdog / restart lists
|
||||
HOST1_OLLAMA_GPU_UUID="GPU-309357d8-2a13-09e0-84ac-fcfdcdf5c626"
|
||||
HOST1_OLLAMA_MODEL="qwen2.5-coder:14b" # generation
|
||||
HOST1_OLLAMA_EMBED_MODEL="nomic-embed-text" # embeddings — qwen cannot embed
|
||||
```
|
||||
|
||||
A node with an empty `OLLAMA_URL` is not an error — it falls through to the resolver and
|
||||
uses the mesh. HOST2 gets the section with blanks, exactly like the NPM/lldap credentials
|
||||
it is already waiting on.
|
||||
|
||||
### `master.conf` — shared behaviour
|
||||
|
||||
```bash
|
||||
# ━━━ AI ━━━
|
||||
AI_ENABLED=false # master switch — fail-closed, != "true" means off
|
||||
AI_CONNECT_TIMEOUT=5 # probe timeout when resolving a host
|
||||
AI_REQUEST_TIMEOUT=240 # must clear a cold load — measured 1m45s after tuning.
|
||||
# KEEP_ALIVE=-1 means this only bites after a restart,
|
||||
# but a first call that times out is the worst first
|
||||
# impression a caller can have. Re-measure against a
|
||||
# real RAG query before fixing this number.
|
||||
AI_RESOLVE_CACHE_TTL=300 # don't re-probe the mesh on every script invocation
|
||||
AI_MAX_RETRIES=1 # AI is enhancement — do not retry hard
|
||||
|
||||
# Per-feature toggles — enable narration long before enabling decision aid
|
||||
AI_ASSIST_REPORTS=false # tier 1 — digest / coffee report narration
|
||||
AI_ASSIST_WATCHDOG=false # tier 2 — context on a flagged condition
|
||||
AI_ASSIST_DISCOVERY=false # tier 2 — discovery / classification judgment calls
|
||||
AI_ASSIST_CLEANUP=false # tier 2 — HELD orphans, stuck-import triage
|
||||
AI_ASSIST_ONBOARD=false # tier 3 — onboarding / settings assistance
|
||||
|
||||
# Conf writes — separate switch, off by default, see Conf Write Access below
|
||||
AI_CONF_WRITE_ENABLED=false
|
||||
AI_CONF_WRITE_KEYS=() # explicit whitelist; never paths or credentials
|
||||
```
|
||||
|
||||
**Per-feature toggles are load-bearing, not decoration.** They are what lets AI narrate the
|
||||
weekly digest for months before it is ever allowed near a cleanup decision. `AI_ENABLED` is
|
||||
necessary but not sufficient — every feature stays individually off until it has earned it.
|
||||
|
||||
> When these land, `Deployment/master.conf.template` and `Deployment/host.conf.template`
|
||||
> must be updated in the same pass. That rule is not optional in this repo.
|
||||
|
||||
---
|
||||
|
||||
## Host resolution
|
||||
|
||||
AI runs on the owner's node only. Remote mesh nodes reach it over Tailscale. No remote node
|
||||
needs a GPU, a model, or an Ollama container — only the resolver.
|
||||
|
||||
`resolve_ollama_host()` mirrors the existing Gitea locator in `git_pull_execute.sh`:
|
||||
|
||||
```
|
||||
Ollama answering on localhost:11434? → use it (owner's node)
|
||||
else discover_remote_nodes()
|
||||
→ resolve_tailscale_ip(node)
|
||||
→ probe each :11434 → first responsive wins
|
||||
else → no AI host (== AI_ENABLED=false)
|
||||
```
|
||||
|
||||
Helpers already exist in `common.sh`: `discover_remote_nodes()` (767),
|
||||
`resolve_tailscale_ip()` (812), `check_connectivity()` (866).
|
||||
|
||||
**Probe the API, not the container.** The Gitea locator checks `docker ps`. Do not copy that
|
||||
here — a container can be up while the model is unloaded, still pulling, or wedged. Probe
|
||||
`/api/tags`. Same principle written into `network_watchdog.sh`'s design principles:
|
||||
*verify the path, not the process*.
|
||||
|
||||
**Cache the resolution** in `/tmp` state, like the arr cache. A cleanup script should not pay
|
||||
a Tailscale round-trip to discover AI it may never call.
|
||||
|
||||
### Known consequence
|
||||
|
||||
AI lives on HOST1, so **during a fallback — HOST1 down, HOST2 covering — the mesh has no AI.**
|
||||
That is precisely when a triage assistant would be most useful. Accepted: a second GPU on
|
||||
HOST2 is a lot of hardware for that window, and AI is enhancement-only by design. Worth
|
||||
knowing rather than discovering.
|
||||
|
||||
---
|
||||
|
||||
## Where AI is allowed to act
|
||||
|
||||
Ranked by how much damage a wrong answer does.
|
||||
|
||||
**Tier 1 — narration and summary (safe, do first)**
|
||||
- Sunday morning coffee report — turn metrics into prose
|
||||
- `weekly_health_digest.sh` — summarise, highlight what changed
|
||||
- Explain *why* a container is crash-looping from its logs
|
||||
|
||||
**Tier 2 — triage and context on an existing flag (the real value)**
|
||||
Places where a script already detects something and stops:
|
||||
- `system_watchdog` / `stability_watchdog` flags a condition → AI adds context, correlates
|
||||
with recent logs, suggests likely cause
|
||||
- Sonarr stuck-import triage — the "matched by series ID" recipe is textbook LLM work
|
||||
- `HELD` entries from `arr_download_orphan_cleaner.sh`
|
||||
- `reverse-anime-leak` from the classification scans — currently report-only *because* it is
|
||||
a judgment call. That is exactly the shape AI suits.
|
||||
|
||||
**Tier 3 — assisted configuration (needs the guardrails below)**
|
||||
- Onboarding a new host — the main motivation for conf write access
|
||||
- AI-assisted settings tuning: rsync profiles, fallback tiers, auth stack
|
||||
|
||||
**Never**
|
||||
- Deciding what to delete
|
||||
- Choosing a path for any destructive operation
|
||||
- Anything that bypasses a strike counter, age gate, or verification step
|
||||
|
||||
---
|
||||
|
||||
## Conf write access
|
||||
|
||||
Wanted mainly for onboarding and assisted settings. This is the highest-risk item here.
|
||||
|
||||
**Current state: there is no recovery path.**
|
||||
|
||||
```
|
||||
.gitignore:4 Configurations/host*.conf
|
||||
.gitignore:5 Configurations/master.conf
|
||||
.gitignore:6 Configurations/*.bak
|
||||
```
|
||||
|
||||
Confs are gitignored — no git history to revert to — and so are the `.bak` files, so the
|
||||
backup is not versioned either. The only fallback is a single `.bak` slot written by
|
||||
`conf_upgrade`, and it goes stale immediately:
|
||||
|
||||
| file | modified | its `.bak` |
|
||||
|---|---|---|
|
||||
| `master.conf` | Jul 28 18:52 | Jul 28 18:52 |
|
||||
| `host1.conf` | Aug 1 21:00 | **Jul 3 17:46** |
|
||||
|
||||
A bad write to `master.conf` currently falls back to a file that may predate a month of edits.
|
||||
**Fix this before any AI writes anything.**
|
||||
|
||||
### Required before conf-write ships
|
||||
|
||||
1. **Key whitelist, not file access.** Thresholds and toggles only — `*_WARN_GB`,
|
||||
`*_STRIKE_LIMIT`, `*_ENABLED`, retention days. Never a path, never a credential, never a
|
||||
container list. A wrong threshold is recoverable; a wrong path is what the depth guards
|
||||
exist to catch.
|
||||
2. **Timestamped backups, plural** — `master.conf.2026-08-01T21:00`, retained. Not one
|
||||
clobbered slot.
|
||||
3. **Validate before commit** — `bash -n` the candidate, then confirm `load_config.sh`
|
||||
sources it cleanly. Never install a conf that has not been proven to parse.
|
||||
4. **Diff always logged.** An AI conf change should be at least as visible as a container
|
||||
restart.
|
||||
5. **Lock against concurrent readers** — never rewrite a conf while scripts are mid-run.
|
||||
|
||||
**Consider un-ignoring `Configurations/` into a private repo.** Then `git diff` and
|
||||
`git revert` become the recovery mechanism and the history is free. This overlaps the
|
||||
existing GitHub-mirror TODO, which is already blocked on the same question.
|
||||
|
||||
---
|
||||
|
||||
## Security
|
||||
|
||||
Ollama has **no authentication of any kind**, and its API includes `DELETE /api/delete`
|
||||
(wipe models) and `POST /api/pull` (fill the disk). It currently binds `0.0.0.0:11434` with
|
||||
`OLLAMA_ORIGINS=*` — reachable from the entire LAN, not just Tailscale.
|
||||
|
||||
The design only needs loopback (owner) plus the Tailscale interface (mesh). `0.0.0.0` is
|
||||
strictly wider than required, for no benefit. Restrict to loopback + Tailscale, or use
|
||||
Tailscale ACLs to allow only mesh nodes. Node-level ACLs fit the mesh model better than
|
||||
app-level auth Ollama cannot provide anyway.
|
||||
|
||||
**Done 2026-08-01:** `/ext-varaverk` is now mounted `ro` (was `rw` into live prod).
|
||||
Verified on the running container — `rw=false`.
|
||||
|
||||
**Still open:** the LAN exposure above. Deliberately not folded into the tuning rebuild,
|
||||
since bind-address versus Tailscale ACL is a decision rather than a setting.
|
||||
|
||||
---
|
||||
|
||||
## Hardware budget
|
||||
|
||||
RTX 3080, 10 GB, pinned to Ollama by UUID — isolated from the Quadro P2000 that Emby
|
||||
transcodes on. Do not let AI onto the P2000.
|
||||
|
||||
**Tuned and measured 2026-08-01.** These are observed values, not estimates.
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| `OLLAMA_NUM_PARALLEL` | 2 | **1** |
|
||||
| `OLLAMA_KV_CACHE_TYPE` | f16 | **q8_0** |
|
||||
| `OLLAMA_CONTEXT_LENGTH` | 4096 | **16384** |
|
||||
| `OLLAMA_FLASH_ATTENTION` | false | **true** |
|
||||
| VRAM used | 9298 MiB (91%) | **8811 MiB (86%)** |
|
||||
| warm latency | 2.6s | **1.85s** |
|
||||
| cold load | 28.6s | **1m45s** |
|
||||
|
||||
`/api/ps` confirms `ctx=16384` — the increase is real, not just an env var.
|
||||
|
||||
**4× the context for less VRAM than before.** Flash Attention plus the quantized KV cache
|
||||
more than paid for the increase. Cold load got much slower, which is irrelevant while
|
||||
`OLLAMA_KEEP_ALIVE=-1` pins both models — but it is felt after any container restart.
|
||||
|
||||
### Flash Attention is mandatory, not optional
|
||||
|
||||
`OLLAMA_KV_CACHE_TYPE=q8_0` **will not load** without it:
|
||||
|
||||
```
|
||||
llama_init_from_model: V cache quantization requires flash_attn
|
||||
llama-server process no longer running: exit status 1
|
||||
```
|
||||
|
||||
Quantized K/V cache requires Flash Attention. Supported on Ampere and newer; the 3080
|
||||
qualifies. If the KV cache type is ever changed back toward a quantized value, Flash
|
||||
Attention must be on or the model silently fails to load and every call errors.
|
||||
|
||||
### Tuning order still matters
|
||||
|
||||
If these are ever re-tuned from defaults, the order is load-bearing — raising context first
|
||||
at high utilisation will OOM:
|
||||
|
||||
1. `OLLAMA_NUM_PARALLEL` → 1 (each slot multiplies KV cache)
|
||||
2. `OLLAMA_FLASH_ATTENTION` → true (prerequisite for the next step)
|
||||
3. `OLLAMA_KV_CACHE_TYPE` → q8_0 (roughly halves KV memory)
|
||||
4. *then* `OLLAMA_CONTEXT_LENGTH` upward
|
||||
|
||||
32k was considered and rejected — projected ~1000 MiB of KV, leaving under 450 MiB headroom.
|
||||
16k is the comfortable ceiling for a 14B on this card.
|
||||
|
||||
### Concurrency
|
||||
|
||||
**You cannot have 14B + long context + real parallelism on 10 GB.** Pick two.
|
||||
`OLLAMA_MAX_QUEUE=512` means excess requests queue rather than fail, and at ~2s responses,
|
||||
two or three users serialised is barely noticeable. Multi-user hits are expected to be rare.
|
||||
So: parallelism stays at 1, the queue absorbs bursts, and the VRAM goes to **context** —
|
||||
which is what RAG actually needs.
|
||||
|
||||
### Applying template changes
|
||||
|
||||
**Unraid's "Apply" does not reliably recreate the container.** Observed 2026-08-01: the
|
||||
template was saved correctly but the container was only *restarted*, so the env vars never
|
||||
took effect — `Created` stayed unchanged while `Started` advanced. Env changes require a
|
||||
remove-and-recreate.
|
||||
|
||||
Force it with Unraid's own script:
|
||||
|
||||
```bash
|
||||
/usr/local/emhttp/plugins/dynamix.docker.manager/scripts/rebuild_container Ollama
|
||||
docker start Ollama # rebuild stops it — Ollama is not in unraid-autostart
|
||||
```
|
||||
|
||||
Verify with `docker inspect Ollama --format '{{.Created}}'` — the timestamp must move.
|
||||
Checking the env vars alone is not enough; a restart leaves the old ones in place and
|
||||
looks like nothing happened.
|
||||
|
||||
---
|
||||
|
||||
## RAG
|
||||
|
||||
No Python on Unraid, and none needed. Everything required is already present:
|
||||
`sqlite3 3.53`, `jq 1.8`, `node 22`, `php 8.4`, `awk`.
|
||||
|
||||
- **chunks + vectors** → SQLite, one table
|
||||
- **similarity** → cosine in PHP or Node; milliseconds at this corpus size, no vector DB
|
||||
container needed
|
||||
- **embed + generate** → Ollama HTTP, same `curl` pattern as every other integration here
|
||||
|
||||
Note `qwen2.5-coder` returns `501 — does not support embeddings`. Embedding is
|
||||
`nomic-embed-text`'s job. Batch embedding works (n inputs → n vectors in one call) and is
|
||||
required — indexing 50k lines one HTTP call at a time is not viable.
|
||||
|
||||
---
|
||||
|
||||
### The corpus — and why its shape matters more than its size
|
||||
|
||||
As of 2026-08-01, after the header audit and the per-folder documentation pass:
|
||||
|
||||
| Layer | Size | What it answers |
|
||||
|-------|------|-----------------|
|
||||
| Script headers | 115 files × 6 sections = **690 chunks**, 14,685 lines | "What does *this script* do, and why that way" |
|
||||
| Folder docs | 18 `README-*.md` + 13 `Manual-*.md` | "How does this *group* work" / "how do I do the thing" |
|
||||
| Top-level | `README.md`, `Manual.md` | "What is this system" |
|
||||
| Conf templates | 2,175 lines, **~55% comment** | The schema, self-describing |
|
||||
| Bash bodies | 51,166 lines | Implementation — index last, lowest weight |
|
||||
|
||||
Markdown total: **13,516 lines.** Still small enough that cosine over the whole set is
|
||||
milliseconds.
|
||||
|
||||
**Chunking is already solved, and the audit is what solved it.** Every script carries
|
||||
`PURPOSE / OPERATIONAL MODEL / DESIGN PRINCIPLES / OPERATIONAL SAFEGUARDS / CONFIGURATION /
|
||||
RUNTIME MODES` — **115 of 115, no exceptions.** Split on `^# SECTION NAME$` and every chunk is
|
||||
a semantically coherent unit by construction. The single worst failure mode in naive RAG —
|
||||
a fixed-size window cutting mid-thought and embedding two half-ideas as one vector — cannot
|
||||
happen here. Median header is 117 lines, so a section lands around 130–200 tokens: comfortably
|
||||
inside `nomic-embed-text`'s window, no sub-splitting needed.
|
||||
|
||||
**Store the section name as a column, not just as chunk text.** This is the highest-value
|
||||
thing the audit bought and it should not be thrown away at index time. Section type is a free
|
||||
metadata filter, so retrieval can route before it computes similarity:
|
||||
|
||||
| Question shape | Filter to |
|
||||
|----------------|-----------|
|
||||
| "what stops X and Y overlapping" | `OPERATIONAL SAFEGUARDS` |
|
||||
| "what variable controls X" | `CONFIGURATION` |
|
||||
| "does this take --dry-run" | `RUNTIME MODES` |
|
||||
| "why is it built this way" | `DESIGN PRINCIPLES` |
|
||||
| "what does this script do" | `PURPOSE` |
|
||||
|
||||
Hybrid retrieval essentially for free, because every chunk already has a type.
|
||||
|
||||
Suggested table shape:
|
||||
|
||||
```sql
|
||||
CREATE TABLE vv_chunks (
|
||||
id INTEGER PRIMARY KEY,
|
||||
path TEXT NOT NULL, -- repo-relative
|
||||
kind TEXT NOT NULL, -- header | readme | manual | template | body
|
||||
section TEXT, -- PURPOSE, OPERATIONAL SAFEGUARDS, ... (NULL for md/body)
|
||||
heading TEXT, -- md ## heading, for doc chunks
|
||||
content TEXT NOT NULL,
|
||||
vector BLOB NOT NULL, -- 768 float32
|
||||
indexed INTEGER NOT NULL -- epoch; re-embed on mtime change only
|
||||
);
|
||||
```
|
||||
|
||||
### Why this corpus is worth more than an equivalent pile of code
|
||||
|
||||
A model can read `mover_stop.sh` and describe what it does. What it *cannot* derive from any
|
||||
amount of source is that a thing was done deliberately. The audit wrote those down:
|
||||
|
||||
- the API cache writers are lockless and unprivileged **on purpose** — regenerable within a
|
||||
minute, every consumer has a live fallback
|
||||
- `removeCompletedDownloads` / `removeFailedDownloads` both true is **intended**, not an
|
||||
oversight
|
||||
- the arr cleanup ctime gate depends on `media_shares_permissions.sh` staying conditional —
|
||||
reverting either silently stops orphan collection
|
||||
- `mesh_monitor.sh`, `adapter.sh`, `decision_engine.sh`, `containers.sh` and
|
||||
`api_cache_writer.sh` carry no root check and no lock **by design** — each documents why in
|
||||
its own header (libraries that must not `exit`, read-only probes, or regenerable output
|
||||
with a live fallback)
|
||||
|
||||
Without those in the index, the most likely contribution from an AI assistant reviewing this
|
||||
repo is a confident regression: *"I notice this script lacks a lock."* Weight
|
||||
`DESIGN PRINCIPLES` and `OPERATIONAL SAFEGUARDS` heavily for any suggest-a-change flow —
|
||||
they are the guardrails against the assistant helpfully undoing a decision.
|
||||
|
||||
### Indexing is safe by default — keep it that way
|
||||
|
||||
`Configurations/*.conf` is gitignored; `Deployment/*.template` is tracked and carries all the
|
||||
explanatory comments. The corpus therefore describes the full schema while structurally
|
||||
**never containing a credential**, because the credential-bearing files were never in the repo
|
||||
to begin with.
|
||||
|
||||
Treat that as a deliberate boundary, not a happy accident:
|
||||
|
||||
- **index tracked files only** — never walk `Configurations/` or `data/`
|
||||
- a live conf value that the model genuinely needs should arrive through a *tool call* at
|
||||
query time, subject to the same redaction rules as everything else in the Security section,
|
||||
not be baked into a vector at index time
|
||||
- an embedded secret is unrevocable in a way a logged one is not — there is no rotation story
|
||||
for a value already averaged into a 768-dim float
|
||||
|
||||
### Known gap — the PHP layer is not covered
|
||||
|
||||
78 PHP files under `Plugin/unraid/`; **2** carry a `PURPOSE` block. The entire web UI —
|
||||
`pages/`, `api/`, `include/` — is effectively invisible to retrieval.
|
||||
|
||||
Consequence: any "AI helper per Varaverk page" feature has this as a hard prerequisite. A
|
||||
page-scoped assistant that cannot retrieve the page's own logic is worse than no assistant.
|
||||
|
||||
`include/` is the high-value subset to do first — 16 files, and both the pages and the API
|
||||
endpoints route through the same `vv_*()` builders, so documenting it once covers both
|
||||
callers. This is a follow-on pass, not a blocker for indexing bash.
|
||||
|
||||
---
|
||||
|
||||
## Scheduled AI
|
||||
|
||||
Same orchestrator tiers as everything else, gated on `AI_ENABLED` plus a reachable host.
|
||||
Natural fits: weekly digest narration, a periodic pass over `HELD`/report-only findings that
|
||||
have accumulated, post-incident summaries after a watchdog event.
|
||||
|
||||
Must obey the existing tier discipline — an AI job that fails or times out is a non-fatal
|
||||
step like any other, and never blocks the rest of its tier.
|
||||
|
||||
---
|
||||
|
||||
## UI
|
||||
|
||||
- **Dedicated AI page** in the plugin.
|
||||
- **Persistent conversation across pages.** A floating widget is not required — a chat column
|
||||
is fine. The constraint is that Unraid's WebGUI is multi-page PHP with full reloads and no
|
||||
SPA shell, so persistence means conversation state lives server-side keyed by session, with
|
||||
the client re-hydrating per page.
|
||||
- **Per-page AI helpers** — contextual assistance scoped to whatever that page is about.
|
||||
|
||||
---
|
||||
|
||||
## "Too bad we can't just run the LLM inside Varaverk and cut out Ollama"
|
||||
|
||||
There is a real answer: you don't cut Ollama out, you **absorb it**.
|
||||
|
||||
Ollama does non-trivial work — model lifecycle, GPU scheduling, keep-alive, batching, an HTTP
|
||||
API. Reimplementing that in bash is not a good trade. But Varaverk already manages containers
|
||||
better than most things manage containers. Ollama becomes just another managed container:
|
||||
|
||||
- add to `HOST*_WATCHDOG_CONTAINERS` so `docker_watchdog.sh` keeps it healthy
|
||||
- add to a restart list so it gets the same proactive treatment as everything else
|
||||
- give it a fallback tier if AI should survive a host outage
|
||||
- let `docker_update.sh` handle its image updates
|
||||
|
||||
That is more Varaverk-native than embedding a model runtime would be, and it costs nothing
|
||||
new — the machinery already exists and was audited this session.
|
||||
|
||||
---
|
||||
|
||||
## Open questions
|
||||
|
||||
- Un-ignore `Configurations/` into a private repo for conf history? (blocks conf-write, and
|
||||
overlaps the GitHub-mirror TODO)
|
||||
- Does the AI page need auth separate from the Unraid WebGUI, given remote mesh members?
|
||||
- Retention/privacy for conversation history — logs may contain paths, container names,
|
||||
possibly credentials pasted by a user.
|
||||
- Is a 7B worth it to buy context + parallelism headroom, or is 14B quality worth the
|
||||
serialisation? Defer until an actual problem is felt.
|
||||
|
||||
---
|
||||
|
||||
## RAG — built 2026-08-02
|
||||
|
||||
Retrieval is live. `AI/` holds the implementation; `AI/README-AI.md` documents it in full. What
|
||||
follows is only what changed relative to the plan recorded above.
|
||||
|
||||
**Corpus is larger than estimated.** ~2,950 chunks across ~180 files, not the 690 header chunks
|
||||
projected. Sub-chunking is why — see below.
|
||||
|
||||
**Named-paragraph sub-chunking was necessary, and was not in the plan.** Section-level chunks
|
||||
alone were too coarse. `rsync.sh` documents fourteen safeguards in one 2.8k-char
|
||||
`OPERATIONAL SAFEGUARDS` block; a query about one of them scored 0.558, below unrelated chunks,
|
||||
because the other thirteen dominated the vector. Splitting on the named-paragraph titles the
|
||||
header convention already uses took the same query to 0.718 and first place. The parent section
|
||||
name is carried onto each sub-chunk so routing still works.
|
||||
|
||||
**Two chunker bugs worth remembering.** The last section in a header (RUNTIME MODES in bash,
|
||||
DEPENDS ON in a page) ran to EOF and swept up every unrelated comment in the file —
|
||||
`scheduler.php` alone produced an 11k-char chunk of unrelated inline comments. And title
|
||||
detection must require the *next* line to be indented; without that, any wrapped prose line
|
||||
became a spurious chunk boundary mid-sentence.
|
||||
|
||||
**Section routing is a boost, not a filter.** Intent detection is a heuristic and must not be
|
||||
able to exclude the chunk holding the answer. `--section=` forces a hard filter when wanted.
|
||||
|
||||
**Vectors arrive pre-normalised.** `nomic-embed-text` returns L2-normalised vectors (measured
|
||||
norm 1.0000001), so cosine is a plain dot product. No normalising step, no magnitude cache.
|
||||
|
||||
**`node:sqlite` over a native module.** Still flagged experimental, chosen because it needs no
|
||||
native compilation on Unraid. Acceptable because the index is disposable — if a Node upgrade
|
||||
breaks it, rebuild takes minutes. PHP reads the same float32 blobs with `unpack('f*', $blob)`
|
||||
when the UI needs them.
|
||||
|
||||
**Retrieval quality, measured.** 9/10 top-3 hit rate on known-answer questions; the tenth had
|
||||
the answer at ranks 2 and 3, so 10/10 for answer-present-in-context at k=6. Full retrieval plus
|
||||
generation runs about 43s warm.
|
||||
|
||||
### The finding that validated the whole thing
|
||||
|
||||
First real end-to-end question asked which variable controls the mover's grace window. The model
|
||||
answered `MOVER_STOP_TIMEOUT`, "defaults to 30 seconds", citing `mover_stop.sh › CONFIGURATION`.
|
||||
Variable correct; the 30 was wrong — the real value is 300. The model was quoting the header
|
||||
verbatim. **The header was stale.**
|
||||
|
||||
A sweep for the same pattern found six stale `(default: N)` claims across the repo — all
|
||||
corrected in the same pass. This is the operating principle for the folder:
|
||||
|
||||
> Retrieval is exactly as accurate as the documentation it points at. When an answer looks
|
||||
> wrong, check the cited source before blaming the model.
|
||||
|
||||
It also means the index is a documentation-drift detector, not only a question-answering tool.
|
||||
|
||||
### Still deliberately not built
|
||||
|
||||
Nothing consults this. Every `AI_ASSIST_*` toggle is false, `AI_CONF_WRITE_ENABLED` is false
|
||||
with an empty key whitelist, and no watchdog, cleanup or fallback path calls it. Host resolution
|
||||
across the Tailscale mesh is specified above but not implemented — `ai_index.sh` and
|
||||
`ai_query.sh` currently require a local `HOST*_OLLAMA_URL` and fail with a clear message when it
|
||||
is empty, rather than silently probing the mesh.
|
||||
@@ -1,47 +0,0 @@
|
||||
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
|
||||
source ~/.bashrc
|
||||
claude
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
❯ now we should be able to allow or have other host be able to start a shared srvice if desired using
|
||||
shared auth stack, maybe a container array in host*.conf folder that can add the container to the other
|
||||
servers, just thinkig of future design
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
. fix failover strike list timing, maybe 30 seconds. them a t 90 seconds 3 stike triggers. just gotta test buffer. never had the strike system
|
||||
. verify silent toggle switches back on good notifications
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
WAY LATER
|
||||
|
||||
now we should be able to allow or have other host be able to start a shared srvice if desired using
|
||||
shared auth stack, maybe a container array in host*.conf folder that can add the container to the other
|
||||
servers, just thinkig of future design
|
||||
|
||||
|
||||
when a owner offboards and there is more than 1 server left, one needs to become the owner, we could do this in multiple ways or a combination, strongest server and BANDWIDTH,whos contributing more.... maybe we just promt and ask them
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
look into an overall, setup script, pull as many vars as possible without user having to add. like docker names, and so on.
|
||||
@@ -1,872 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
# ==============================================================================================
|
||||
# ================================= Arr Cleanup ================================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Delete orphaned media files not tracked by an arr (Lidarr, Radarr, Sonarr,
|
||||
# or any future arr). Queries the API for all tracked file paths, walks the
|
||||
# library on disk, and removes anything untracked that is old enough to be
|
||||
# past the import window. Triggers an Emby library clean after each deletion
|
||||
# run so ghost entries disappear immediately.
|
||||
#
|
||||
# Called exclusively by arr_cleanup.sh, which sources shell config and exports
|
||||
# all configuration as environment variables before exec'ing this script.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Every file encountered on disk is classified into one of five categories:
|
||||
#
|
||||
# TRACKED — arr API knows this exact path → leave it alone
|
||||
# PROTECTED — matches {ARR}_PROTECTED_PATTERNS → never delete
|
||||
# ORPHAN — media file, not tracked, older than {ARR}_ORPHAN_AGE → delete
|
||||
# JUNK — not a tracked extension, not protected → delete regardless of age
|
||||
# RECENT — not tracked, under {ARR}_ORPHAN_AGE → skip (may be mid-import)
|
||||
#
|
||||
# Arr apps generate cover art (*.jpg), metadata (*.nfo), lyrics (*.lrc), and
|
||||
# subtitles (*.srt) but do NOT include these in their tracked file API response.
|
||||
# Without PROTECTED classification these would be deleted — breaking the arr
|
||||
# app and Emby display.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Seven gates — ALL must pass before any file is touched:
|
||||
# 1. Container running and not starting/unhealthy
|
||||
# 2. API reachable
|
||||
# 3. API version matches {ARR}_VERSION_MAJOR in master.conf
|
||||
# 4. Parent count > 0 (artists / movies / series)
|
||||
# 5. Tracked file count > 0
|
||||
# 6. Tracked count >= {ARR}_MIN_TRACKED_PCT % of last known (if configured)
|
||||
# 7. Deletion size < {ARR}_MAX_DELETE_GB — or --i-know-what-im-doing required
|
||||
#
|
||||
# ==============================================================================================
|
||||
# ADDING A NEW ARR
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. Add an entry to ARR_PROFILES below (6 values — API endpoint pattern only)
|
||||
# 2. Add HOST*_<ARR>_URL, API_KEY, MEDIA_ROOT, PATH_MAP to master_host*.conf
|
||||
# 3. Add <ARR>_ORPHAN_AGE, MAX_DELETE_GB, EXTENSIONS, etc. to master.conf
|
||||
# 4. Add an export block to arr_cleanup.sh (copy existing block, change prefix)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master_host*.conf (host-specific, aliased by detect_hosts in arr_cleanup.sh)
|
||||
#
|
||||
# HOST*_{ARR}_URL — arr base URL
|
||||
# HOST*_{ARR}_API_KEY — arr API key
|
||||
# HOST*_{ARR}_MEDIA_ROOT — host-side library root (MUSIC_ROOT / MOVIES_ROOT / TV_ROOT)
|
||||
# HOST*_{ARR}_PATH_MAP — container path → host path translation (assoc array)
|
||||
#
|
||||
# master.conf (shared thresholds)
|
||||
#
|
||||
# {ARR}_ORPHAN_AGE — days before untracked file is eligible for deletion
|
||||
# {ARR}_MAX_DELETE_GB — require --i-know-what-im-doing above this
|
||||
# {ARR}_MIN_TRACKED_PCT — abort if tracked count drops below this % of last run
|
||||
# {ARR}_TRACKED_COUNT_FILE — persistent baseline file path (enables gate 6)
|
||||
# {ARR}_EXTENSIONS — media file extensions for orphan classification
|
||||
# {ARR}_PROTECTED_PATTERNS — glob patterns never deleted (cover art, metadata, etc.)
|
||||
# {ARR}_VERSION_MAJOR — expected arr major version for API safety check
|
||||
# {ARR}_IMPORT_SCAN_TIMEOUT — seconds to wait for pre-flight import scan (default 600)
|
||||
# {ARR}_LOCK_WARN_AGE — override default lock warning age (large libraries)
|
||||
# ARR_CLEANUP_STATS — stats file path (read by sunday_morning_coffee_report)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# arr_cleanup.sh --arr lidarr — normal run
|
||||
# arr_cleanup.sh --arr lidarr --dry-run — preview, no deletions
|
||||
# arr_cleanup.sh --arr lidarr --log — verbose output
|
||||
# arr_cleanup.sh --arr lidarr --status — show config and exit
|
||||
# arr_cleanup.sh --arr lidarr --i-know-what-im-doing — bypass size threshold
|
||||
# arr_cleanup.sh --arr lidarr --i-know-what-im-doing --skip-strike-list — NUCLEAR MODE
|
||||
#
|
||||
# NUCLEAR MODE: both flags bypass age check AND size threshold. Use when the arr
|
||||
# has filled gaps and you want a clean one-pass wipe. Flag name is long and
|
||||
# annoying by design.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
import argparse
|
||||
import atexit
|
||||
import datetime
|
||||
import fnmatch
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
# ── Arr profiles — only what cannot come from env vars ────────────────────────
|
||||
# API endpoint patterns and labels. Everything else is config in master.conf.
|
||||
ARR_PROFILES = {
|
||||
"lidarr": {
|
||||
"api_version": "v1",
|
||||
"parent_endpoint": "artist",
|
||||
"parent_id_param": "artistId",
|
||||
"file_endpoint": "trackFile",
|
||||
"import_scan_cmd": "DownloadedAlbumsScan",
|
||||
"parent_label": "artists",
|
||||
},
|
||||
"radarr": {
|
||||
"api_version": "v3",
|
||||
"parent_endpoint": "movie",
|
||||
"parent_id_param": "movieId",
|
||||
"file_endpoint": "moviefile",
|
||||
"import_scan_cmd": "DownloadedMoviesScan",
|
||||
"parent_label": "movies",
|
||||
},
|
||||
"sonarr": {
|
||||
"api_version": "v3",
|
||||
"parent_endpoint": "series",
|
||||
"parent_id_param": "seriesId",
|
||||
"file_endpoint": "episodefile",
|
||||
"import_scan_cmd": "DownloadedEpisodesScan",
|
||||
"parent_label": "series",
|
||||
},
|
||||
# Add new arrs here. 6 values — everything else goes in master.conf.
|
||||
# "readarr": {
|
||||
# "api_version": "v1",
|
||||
# "parent_endpoint": "author",
|
||||
# "parent_id_param": "authorId",
|
||||
# "file_endpoint": "bookfile",
|
||||
# "import_scan_cmd": "DownloadedBooksScan",
|
||||
# "parent_label": "authors",
|
||||
# },
|
||||
}
|
||||
|
||||
# ── Output helpers ─────────────────────────────────────────────────────────────
|
||||
|
||||
VERBOSE = False
|
||||
|
||||
def log(msg):
|
||||
if VERBOSE:
|
||||
print(f" {msg}")
|
||||
|
||||
def warn(msg):
|
||||
print(f" ⚠️ {msg}")
|
||||
|
||||
def error(msg):
|
||||
print(f" ❌ {msg}", file=sys.stderr)
|
||||
|
||||
def success(msg):
|
||||
print(f" ✅ {msg}")
|
||||
|
||||
def die(msg, notify_fn=None):
|
||||
error(msg)
|
||||
if notify_fn:
|
||||
notify_fn(msg)
|
||||
sys.exit(1)
|
||||
|
||||
# ── Env var readers ────────────────────────────────────────────────────────────
|
||||
|
||||
def _env(key, default=""):
|
||||
return os.environ.get(key, default)
|
||||
|
||||
def _env_int(key, default=0):
|
||||
try:
|
||||
return int(os.environ.get(key, str(default)))
|
||||
except ValueError:
|
||||
return default
|
||||
|
||||
def _env_float(key, default=0.0):
|
||||
try:
|
||||
return float(os.environ.get(key, str(default)))
|
||||
except ValueError:
|
||||
return default
|
||||
|
||||
def _env_list(key):
|
||||
val = os.environ.get(key, "")
|
||||
return val.split() if val else []
|
||||
|
||||
def _env_bool(key):
|
||||
return os.environ.get(key, "false").lower() == "true"
|
||||
|
||||
# ── Lock ───────────────────────────────────────────────────────────────────────
|
||||
|
||||
LOCK_DIR = _env("LOCK_DIR", "/tmp/unraid_locks")
|
||||
LOCK_TIMEOUT = _env_int("LOCK_WAIT_TIMEOUT", 30)
|
||||
SCRIPT_NAME = "arr_cleanup"
|
||||
|
||||
def acquire_lock(warn_age=3600):
|
||||
os.makedirs(LOCK_DIR, exist_ok=True)
|
||||
lockfile = Path(LOCK_DIR) / f"{SCRIPT_NAME}.lock"
|
||||
|
||||
if lockfile.exists():
|
||||
try:
|
||||
content = lockfile.read_text().strip()
|
||||
pid_str, locked_name = content.split(":", 1)
|
||||
pid = int(pid_str)
|
||||
|
||||
try:
|
||||
os.kill(pid, 0)
|
||||
pid_alive = True
|
||||
except (ProcessLookupError, PermissionError):
|
||||
pid_alive = False
|
||||
|
||||
if not pid_alive or locked_name != SCRIPT_NAME:
|
||||
warn(f"Stale lock (PID {pid} gone) — clearing")
|
||||
lockfile.unlink(missing_ok=True)
|
||||
else:
|
||||
age = time.time() - lockfile.stat().st_mtime
|
||||
if age > warn_age:
|
||||
warn(f"{SCRIPT_NAME} has been running for {int(age)}s — may be stuck (PID {pid})")
|
||||
print(f" Another instance of {SCRIPT_NAME} is running — waiting up to {LOCK_TIMEOUT}s...")
|
||||
waited = 0
|
||||
while lockfile.exists() and waited < LOCK_TIMEOUT:
|
||||
time.sleep(1)
|
||||
waited += 1
|
||||
if lockfile.exists():
|
||||
die(f"{SCRIPT_NAME} still locked after {LOCK_TIMEOUT}s — exiting")
|
||||
except (ValueError, OSError):
|
||||
lockfile.unlink(missing_ok=True)
|
||||
|
||||
lockfile.write_text(f"{os.getpid()}:{SCRIPT_NAME}")
|
||||
atexit.register(lambda: lockfile.unlink(missing_ok=True))
|
||||
log(f"🔏 Lock acquired: {SCRIPT_NAME} (PID {os.getpid()})")
|
||||
|
||||
# ── HTTP helper ────────────────────────────────────────────────────────────────
|
||||
|
||||
def _http_get(url, api_key, timeout=30):
|
||||
req = urllib.request.Request(url, headers={"X-Api-Key": api_key})
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||
return resp.status, resp.read().decode()
|
||||
except urllib.error.HTTPError as e:
|
||||
return e.code, ""
|
||||
except Exception:
|
||||
return 0, ""
|
||||
|
||||
def _http_post(url, api_key, payload=None, timeout=30):
|
||||
data = json.dumps(payload or {}).encode()
|
||||
req = urllib.request.Request(
|
||||
url, data=data, method="POST",
|
||||
headers={"X-Api-Key": api_key, "Content-Type": "application/json"},
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=timeout) as resp:
|
||||
return resp.status, resp.read().decode()
|
||||
except urllib.error.HTTPError as e:
|
||||
return e.code, ""
|
||||
except Exception:
|
||||
return 0, ""
|
||||
|
||||
# ── Arr API ────────────────────────────────────────────────────────────────────
|
||||
|
||||
def arr_get(base_url, api_key, api_version, endpoint):
|
||||
url = f"{base_url}/api/{api_version}/{endpoint}"
|
||||
status, body = _http_get(url, api_key)
|
||||
if status != 200:
|
||||
error(f"API HTTP {status} for: {endpoint}")
|
||||
return None
|
||||
try:
|
||||
return json.loads(body)
|
||||
except json.JSONDecodeError:
|
||||
error(f"Failed to parse JSON for: {endpoint}")
|
||||
return None
|
||||
|
||||
# ── Path translation ───────────────────────────────────────────────────────────
|
||||
# Replicates common.sh translate_path() — longest-prefix match wins.
|
||||
|
||||
def translate_path(api_path, path_map):
|
||||
best_match = ""
|
||||
best_len = 0
|
||||
for container_path, host_path in path_map.items():
|
||||
if api_path.startswith(container_path) and len(container_path) > best_len:
|
||||
best_match = container_path
|
||||
best_len = len(container_path)
|
||||
if best_match:
|
||||
return path_map[best_match] + api_path[best_len:]
|
||||
return api_path
|
||||
|
||||
# ── Container health check ─────────────────────────────────────────────────────
|
||||
|
||||
def check_container(name, timeout=15):
|
||||
def _inspect(fmt):
|
||||
try:
|
||||
r = subprocess.run(
|
||||
["timeout", str(timeout), "docker", "inspect", "-f", fmt, name],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
return r.stdout.strip()
|
||||
except Exception:
|
||||
return ""
|
||||
|
||||
if _inspect("{{.State.Running}}") != "true":
|
||||
return False, f"{name} is not running"
|
||||
|
||||
health = _inspect("{{.State.Health.Status}}")
|
||||
if health == "healthy":
|
||||
log(f"{name} is healthy")
|
||||
elif health == "":
|
||||
log(f"{name} has no health check — proceeding")
|
||||
elif health == "starting":
|
||||
return False, f"{name} is still starting"
|
||||
elif health == "unhealthy":
|
||||
return False, f"{name} is unhealthy"
|
||||
else:
|
||||
warn(f"{name} health: {health} — proceeding with caution")
|
||||
|
||||
return True, ""
|
||||
|
||||
# ── API reachability ───────────────────────────────────────────────────────────
|
||||
|
||||
def check_api(url, label, timeout=10):
|
||||
req = urllib.request.Request(url)
|
||||
try:
|
||||
urllib.request.urlopen(req, timeout=timeout)
|
||||
log(f"{label} API reachable: {url}")
|
||||
return True
|
||||
except Exception:
|
||||
error(f"{label} API not reachable: {url}")
|
||||
return False
|
||||
|
||||
# ── API version check ──────────────────────────────────────────────────────────
|
||||
|
||||
def check_arr_version(base_url, api_key, api_version, expected_major, label):
|
||||
url = f"{base_url}/api/{api_version}/system/status"
|
||||
status, body = _http_get(url, api_key, timeout=10)
|
||||
if status != 200 or not body:
|
||||
warn(f"{label} version check failed — proceeding without verification")
|
||||
return True
|
||||
try:
|
||||
data = json.loads(body)
|
||||
version = data.get("version", "")
|
||||
major = version.split(".")[0]
|
||||
if major == str(expected_major):
|
||||
success(f"{label} version: {version} (major {major} — tested ✅)")
|
||||
return True
|
||||
else:
|
||||
error(f"{label} version mismatch — running v{major}, tested against v{expected_major}")
|
||||
error(f"API structure may have changed — update {label.upper()}_VERSION_MAJOR in master.conf after verifying")
|
||||
return False
|
||||
except Exception:
|
||||
warn(f"{label} version check failed — could not parse response")
|
||||
return True
|
||||
|
||||
# ── Import scan (pre-flight) ───────────────────────────────────────────────────
|
||||
|
||||
def run_import_scan(base_url, api_key, api_version, scan_cmd, media_root, path_map, timeout=600, label="Arr"):
|
||||
container_root = next(
|
||||
(cp for cp, hp in path_map.items() if hp == media_root),
|
||||
None,
|
||||
)
|
||||
|
||||
if container_root:
|
||||
log(f"Triggering {scan_cmd} on: {container_root}")
|
||||
payload = {"name": scan_cmd, "path": container_root}
|
||||
else:
|
||||
log(f"No path map match — triggering {scan_cmd} (all root folders)")
|
||||
payload = {"name": scan_cmd}
|
||||
|
||||
url = f"{base_url}/api/{api_version}/command"
|
||||
status, body = _http_post(url, api_key, payload)
|
||||
if status not in (200, 201):
|
||||
warn(f"Could not trigger import scan (HTTP {status}) — proceeding without pre-flight")
|
||||
return
|
||||
|
||||
try:
|
||||
cmd_id = json.loads(body).get("id")
|
||||
except Exception:
|
||||
cmd_id = None
|
||||
|
||||
if not cmd_id:
|
||||
warn("Could not get scan command ID — proceeding without pre-flight")
|
||||
return
|
||||
|
||||
print(f" Import scan queued (command ID: {cmd_id}) — waiting for completion...")
|
||||
polled = 0
|
||||
poll_url = f"{base_url}/api/{api_version}/command/{cmd_id}"
|
||||
while polled < timeout:
|
||||
_, body = _http_get(poll_url, api_key, timeout=10)
|
||||
try:
|
||||
state = json.loads(body).get("status", "")
|
||||
except Exception:
|
||||
state = ""
|
||||
|
||||
if state == "completed":
|
||||
log("Import scan complete ✅")
|
||||
return
|
||||
if state == "failed":
|
||||
warn("Import scan reported failed — proceeding anyway")
|
||||
return
|
||||
|
||||
time.sleep(10)
|
||||
polled += 10
|
||||
if polled % 60 == 0:
|
||||
log(f" Still scanning... ({polled}s elapsed)")
|
||||
|
||||
warn(f"Import scan timed out after {timeout}s — proceeding anyway")
|
||||
|
||||
# ── Fetch tracked paths ────────────────────────────────────────────────────────
|
||||
|
||||
def fetch_tracked_paths(base_url, api_key, api_version, profile, path_map, label):
|
||||
parent_endpoint = profile["parent_endpoint"]
|
||||
parent_id_param = profile["parent_id_param"]
|
||||
file_endpoint = profile["file_endpoint"]
|
||||
parent_label = profile["parent_label"]
|
||||
|
||||
parents = arr_get(base_url, api_key, api_version, parent_endpoint)
|
||||
if parents is None:
|
||||
return None, 0, 0
|
||||
|
||||
parent_ids = [p["id"] for p in parents]
|
||||
parent_count = len(parent_ids)
|
||||
if parent_count == 0:
|
||||
return None, 0, 0
|
||||
|
||||
log(f"Found {parent_count} {parent_label} — fetching tracked files...")
|
||||
|
||||
tracked = set()
|
||||
for i, pid in enumerate(parent_ids):
|
||||
if i > 0 and i % 100 == 0:
|
||||
log(f"Fetching files: {i}/{parent_count} {parent_label}...")
|
||||
files = arr_get(base_url, api_key, api_version, f"{file_endpoint}?{parent_id_param}={pid}")
|
||||
if not files:
|
||||
continue
|
||||
if isinstance(files, dict):
|
||||
files = [files]
|
||||
for f in files:
|
||||
api_path = f.get("path", "")
|
||||
if api_path:
|
||||
tracked.add(translate_path(api_path, path_map))
|
||||
|
||||
return tracked, parent_count, len(tracked)
|
||||
|
||||
# ── File classification helpers ────────────────────────────────────────────────
|
||||
|
||||
def is_media_file(path, extensions):
|
||||
ext = Path(path).suffix.lstrip(".").lower()
|
||||
return ext in extensions
|
||||
|
||||
def is_protected(path, patterns):
|
||||
name = Path(path).name
|
||||
return any(fnmatch.fnmatch(name, p) for p in patterns)
|
||||
|
||||
# ── Notify Emby ───────────────────────────────────────────────────────────────
|
||||
|
||||
def notify_emby_scan(emby_url, emby_api_key, my_id):
|
||||
if not emby_url or not emby_api_key:
|
||||
log(f"Emby not configured on {my_id} — skipping library scan notification")
|
||||
return
|
||||
|
||||
log("Notifying Emby to clean missing files...")
|
||||
status, body = _http_get(f"{emby_url}/ScheduledTasks", emby_api_key, timeout=15)
|
||||
if status != 200 or not body:
|
||||
warn(f"Could not reach Emby scheduled tasks API — skipping scan")
|
||||
return
|
||||
|
||||
try:
|
||||
tasks = json.loads(body)
|
||||
except Exception:
|
||||
warn("Could not parse Emby tasks response")
|
||||
return
|
||||
|
||||
task_id = None
|
||||
for task in tasks:
|
||||
if "Clean Missing" in task.get("Name", ""):
|
||||
task_id = task.get("Id")
|
||||
break
|
||||
if not task_id:
|
||||
for task in tasks:
|
||||
if "Scan Media Library" in task.get("Name", ""):
|
||||
task_id = task.get("Id")
|
||||
log("Clean Missing Files not found — using Scan Media Library")
|
||||
break
|
||||
|
||||
if not task_id:
|
||||
warn("Could not find Emby Clean Missing Files or Scan Media Library task")
|
||||
warn("Ghost entries will persist until next Emby scan")
|
||||
return
|
||||
|
||||
status, _ = _http_post(f"{emby_url}/ScheduledTasks/Running/{task_id}", emby_api_key)
|
||||
if status in (200, 204):
|
||||
warn("🎬 Emby Clean Missing Files triggered — ghost entries will be removed")
|
||||
else:
|
||||
warn(f"Emby task trigger returned HTTP {status} — ghost entries may persist")
|
||||
|
||||
# ── Unraid + Discord notifications ────────────────────────────────────────────
|
||||
|
||||
def notify(msg, subject, notify_unraid, hostname, discord_webhook):
|
||||
log(f"🔔 Notification: {subject} — {msg}")
|
||||
if notify_unraid:
|
||||
notify_script = "/usr/local/emhttp/plugins/dynamix/scripts/notify"
|
||||
if os.path.isfile(notify_script) and os.access(notify_script, os.X_OK):
|
||||
subprocess.run([notify_script, "-s", subject, "-d", msg, "-i", "warning"],
|
||||
capture_output=True)
|
||||
if discord_webhook:
|
||||
payload = json.dumps({"content": f"🔔 **{subject}**\n{msg}"}).encode()
|
||||
req = urllib.request.Request(
|
||||
discord_webhook, data=payload, method="POST",
|
||||
headers={"Content-Type": "application/json"},
|
||||
)
|
||||
try:
|
||||
urllib.request.urlopen(req, timeout=10)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
# ── Format helpers ─────────────────────────────────────────────────────────────
|
||||
|
||||
def format_bytes(b):
|
||||
if b > 1_073_741_824:
|
||||
return f"{b / 1_073_741_824:.1f}GB"
|
||||
if b > 1_048_576:
|
||||
return f"{b / 1_048_576:.1f}MB"
|
||||
return f"{b}B"
|
||||
|
||||
def format_duration(secs):
|
||||
if secs >= 3600:
|
||||
return f"{secs // 3600}h{(secs % 3600) // 60}m{secs % 60}s"
|
||||
if secs >= 60:
|
||||
return f"{secs // 60}m{secs % 60}s"
|
||||
return f"{secs}s"
|
||||
|
||||
# ══════════════════════════════════════════════════════════════════════════════
|
||||
# MAIN
|
||||
# ══════════════════════════════════════════════════════════════════════════════
|
||||
|
||||
def main():
|
||||
global VERBOSE
|
||||
|
||||
parser = argparse.ArgumentParser(prog="arr_cleanup.py")
|
||||
parser.add_argument("--arr", required=True, choices=list(ARR_PROFILES.keys()))
|
||||
parser.add_argument("--dry-run", action="store_true")
|
||||
parser.add_argument("--log", action="store_true")
|
||||
parser.add_argument("--status", action="store_true")
|
||||
parser.add_argument("--i-know-what-im-doing", action="store_true", dest="i_know")
|
||||
parser.add_argument("--skip-strike-list", action="store_true", dest="skip_strikes")
|
||||
args = parser.parse_args()
|
||||
|
||||
VERBOSE = args.log
|
||||
ARR = args.arr
|
||||
PREFIX = ARR.upper()
|
||||
profile = ARR_PROFILES[ARR]
|
||||
arr_label = ARR.capitalize()
|
||||
|
||||
# ── Read config from env ───────────────────────────────────────────────────
|
||||
url = _env(f"{PREFIX}_URL")
|
||||
api_key = _env(f"{PREFIX}_API_KEY")
|
||||
media_root = _env(f"{PREFIX}_MEDIA_ROOT")
|
||||
extensions = set(_env_list(f"{PREFIX}_EXTENSIONS"))
|
||||
protected = _env_list(f"{PREFIX}_PROTECTED_PATTERNS")
|
||||
orphan_age = _env_int(f"{PREFIX}_ORPHAN_AGE", 3)
|
||||
max_del_gb = _env_float(f"{PREFIX}_MAX_DELETE_GB", 10.0)
|
||||
min_pct = _env_int(f"{PREFIX}_MIN_TRACKED_PCT", 0)
|
||||
count_file = _env(f"{PREFIX}_TRACKED_COUNT_FILE")
|
||||
ver_major = _env(f"{PREFIX}_VERSION_MAJOR", "0")
|
||||
scan_tmout = _env_int(f"{PREFIX}_IMPORT_SCAN_TIMEOUT", 600)
|
||||
lock_warn = _env_int(f"{PREFIX}_LOCK_WARN_AGE", 3600)
|
||||
container = _env(f"{PREFIX}_CONTAINER") or arr_label
|
||||
|
||||
path_map_json = _env(f"{PREFIX}_PATH_MAP_JSON", "{}")
|
||||
try:
|
||||
path_map = json.loads(path_map_json)
|
||||
except json.JSONDecodeError:
|
||||
path_map = {}
|
||||
|
||||
my_id = _env("MY_ID", "HOST1")
|
||||
server_name = _env("LOCAL_SERVER_NAME")
|
||||
notify_unraid = _env_bool("NOTIFY_UNRAID")
|
||||
emby_url = _env("EMBY_URL")
|
||||
emby_api_key = _env("EMBY_API_KEY")
|
||||
discord = _env("MY_DISCORD_WEBHOOK")
|
||||
stats_file = _env("ARR_CLEANUP_STATS")
|
||||
hostname = _env("ARR_HOSTNAME")
|
||||
api_version = profile["api_version"]
|
||||
|
||||
def _notify(msg, subject=f"{arr_label} Cleanup"):
|
||||
notify(msg, subject, notify_unraid, hostname, discord)
|
||||
|
||||
# ── Nuclear mode warning ───────────────────────────────────────────────────
|
||||
if args.i_know and args.skip_strikes and not args.dry_run:
|
||||
print()
|
||||
print("━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━")
|
||||
print("⚠️ WARNING — NUCLEAR MODE ACTIVE")
|
||||
print("━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━")
|
||||
print(" Flags: --i-know-what-im-doing --skip-strike-list")
|
||||
print(" Strike system: BYPASSED — deletes on first pass")
|
||||
print(" Size threshold: BYPASSED — no GB limit")
|
||||
print(" Data recovery: NOT POSSIBLE after deletion")
|
||||
print()
|
||||
print(" Review --dry-run output before proceeding.")
|
||||
print(" You have 10 seconds to cancel (Ctrl+C)...")
|
||||
print("━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━")
|
||||
time.sleep(10)
|
||||
print(" Proceeding...")
|
||||
print()
|
||||
|
||||
# ── Setup ──────────────────────────────────────────────────────────────────
|
||||
print()
|
||||
print("━━━ ⚙️ Setup ━━━")
|
||||
|
||||
if not url or not api_key:
|
||||
log(f"{arr_label} not configured on {my_id} ({server_name}) — skipping")
|
||||
sys.exit(0)
|
||||
|
||||
if not media_root:
|
||||
die(f"{PREFIX}_MEDIA_ROOT not set — check master_host*.conf", _notify)
|
||||
|
||||
if not Path(media_root).is_dir():
|
||||
die(f"Media root not found: {media_root}", _notify)
|
||||
|
||||
acquire_lock(warn_age=lock_warn)
|
||||
|
||||
print(f" {my_id} ({server_name}) — {url}")
|
||||
|
||||
if args.dry_run: warn("DRY RUN — no files will be deleted")
|
||||
if args.i_know: warn("OVERRIDE — --i-know-what-im-doing active")
|
||||
if args.skip_strikes: warn("OVERRIDE — --skip-strike-list active — age check bypassed")
|
||||
|
||||
# ── Status ─────────────────────────────────────────────────────────────────
|
||||
if args.status:
|
||||
print()
|
||||
print(f"━━━━━ 📋 STATUS ━━━━━")
|
||||
print(f"⚙️ Identity: {my_id} ({server_name})")
|
||||
print(f"⚙️ {arr_label} URL: {url}")
|
||||
print(f"⚙️ Media root: {media_root}")
|
||||
print(f"⏱️ Orphan age: {orphan_age} days")
|
||||
print(f"⚙️ Max delete: {max_del_gb}GB (requires --i-know-what-im-doing)")
|
||||
if min_pct:
|
||||
print(f"⚙️ Min tracked %: {min_pct}%")
|
||||
print(f"⚙️ {arr_label} ver: v{ver_major} expected")
|
||||
print(f"⚙️ Extensions: {' '.join(sorted(extensions))}")
|
||||
print(f"⚙️ Protected patterns: {' '.join(protected)}")
|
||||
print(f"⚙️ Dry Run: {args.dry_run}")
|
||||
print(f"⚙️ I know: {args.i_know}")
|
||||
print(f"⚙️ Skip strikes: {args.skip_strikes}")
|
||||
print("━━━━━━━━━━━━━━━━━━━━━━━")
|
||||
sys.exit(0)
|
||||
|
||||
# ── Safety gate 1 — container health ──────────────────────────────────────
|
||||
print()
|
||||
print("━━━ 🛡️ Safety Checks ━━━")
|
||||
|
||||
ok, reason = check_container(container)
|
||||
if not ok:
|
||||
_notify(f"{arr_label} cleanup aborted on {hostname} — {reason}")
|
||||
die(f"{reason} — aborting")
|
||||
|
||||
log("Safety gate 1 passed — container healthy")
|
||||
|
||||
# ── Pre-flight import scan ─────────────────────────────────────────────────
|
||||
print()
|
||||
print(f"━━━ 🔄 Pre-flight: {arr_label} Import Scan ━━━")
|
||||
run_import_scan(url, api_key, api_version,
|
||||
profile["import_scan_cmd"], media_root, path_map,
|
||||
timeout=scan_tmout, label=arr_label)
|
||||
|
||||
# ── Safety gate 2 — API reachability ──────────────────────────────────────
|
||||
print()
|
||||
print(f"━━━ 🔄 Fetching {arr_label} Tracked Files ━━━")
|
||||
|
||||
if not check_api(url, arr_label):
|
||||
_notify(f"{arr_label} cleanup aborted on {hostname} — API unreachable")
|
||||
sys.exit(1)
|
||||
|
||||
# ── Safety gate 3 — API version ───────────────────────────────────────────
|
||||
if not check_arr_version(url, api_key, api_version, ver_major, arr_label):
|
||||
_notify(f"{arr_label} version mismatch on {hostname} — check master.conf")
|
||||
sys.exit(1)
|
||||
|
||||
# ── Fetch tracked paths ────────────────────────────────────────────────────
|
||||
print(f" Querying {arr_label} API...")
|
||||
tracked, parent_count, tracked_count = fetch_tracked_paths(
|
||||
url, api_key, api_version, profile, path_map, arr_label,
|
||||
)
|
||||
|
||||
# ── Safety gate 4 — parent count > 0 ──────────────────────────────────────
|
||||
if tracked is None or parent_count == 0:
|
||||
msg = f"API returned 0 {profile['parent_label']} — aborting to prevent mass deletion"
|
||||
_notify(f"{arr_label} cleanup aborted on {hostname} — 0 {profile['parent_label']} returned")
|
||||
die(msg)
|
||||
|
||||
# ── Safety gate 5 — tracked count > 0 ────────────────────────────────────
|
||||
if tracked_count == 0:
|
||||
_notify(f"{arr_label} cleanup aborted on {hostname} — 0 tracked files returned")
|
||||
die("API returned 0 tracked files — aborting to prevent mass deletion")
|
||||
|
||||
print(f" {parent_count} {profile['parent_label']} | {tracked_count} tracked files")
|
||||
log(f"Built in-memory lookup set: {tracked_count} tracked paths")
|
||||
|
||||
# ── Safety gate 6 — tracked % drop (only if configured) ──────────────────
|
||||
if count_file and min_pct > 0:
|
||||
count_path = Path(count_file)
|
||||
if count_path.exists():
|
||||
try:
|
||||
last = int(count_path.read_text().strip())
|
||||
if last > 0:
|
||||
pct = int((tracked_count / last) * 100)
|
||||
if pct < min_pct:
|
||||
error(f"Tracked count dropped to {pct}% of last run ({tracked_count} vs {last})")
|
||||
error(f"Suggests API issue — aborting to prevent mass deletion")
|
||||
error(f"If expected (large removal) delete: {count_file}")
|
||||
_notify(f"{arr_label} cleanup aborted on {hostname} — tracked count dropped to {pct}%")
|
||||
sys.exit(1)
|
||||
log(f"Tracked count: {pct}% of last run ({tracked_count} vs {last}) ✅")
|
||||
except (ValueError, OSError):
|
||||
log("Could not read previous count — skipping % check")
|
||||
else:
|
||||
log("No previous count on record — first run, saving baseline")
|
||||
try:
|
||||
count_path.write_text(str(tracked_count))
|
||||
except OSError as e:
|
||||
warn(f"Could not write tracked count file: {e}")
|
||||
|
||||
# ── Scan media root ────────────────────────────────────────────────────────
|
||||
print()
|
||||
print(f"━━━ 🧹 Scanning Media Root ━━━")
|
||||
print(f" Root: {media_root} | Orphan age: {orphan_age} days")
|
||||
print()
|
||||
|
||||
start = time.time()
|
||||
orphan_count = junk_count = recent_count = protected_count = 0
|
||||
orphan_bytes = junk_bytes = 0
|
||||
age_threshold = orphan_age * 86400
|
||||
now = time.time()
|
||||
max_del_bytes = int(max_del_gb * 1_073_741_824)
|
||||
|
||||
scan_roots = set(path_map.values()) | {media_root}
|
||||
|
||||
all_files = []
|
||||
for root in scan_roots:
|
||||
if Path(root).is_dir():
|
||||
for fp in Path(root).rglob("*"):
|
||||
if fp.is_file():
|
||||
all_files.append(str(fp))
|
||||
all_files = sorted(set(all_files))
|
||||
|
||||
for filepath in all_files:
|
||||
if filepath in tracked:
|
||||
log(f"TRACKED: {filepath}")
|
||||
continue
|
||||
|
||||
if is_protected(filepath, protected):
|
||||
log(f"🔰 PROTECTED: {filepath}")
|
||||
protected_count += 1
|
||||
continue
|
||||
|
||||
try:
|
||||
st = Path(filepath).stat()
|
||||
except OSError:
|
||||
continue
|
||||
|
||||
file_size = st.st_size
|
||||
|
||||
if is_media_file(filepath, extensions):
|
||||
file_age = now - st.st_mtime
|
||||
if file_age < age_threshold and not args.skip_strikes:
|
||||
log(f"RECENT (skipping): {filepath}")
|
||||
recent_count += 1
|
||||
continue
|
||||
warn(f"🗑️ ORPHAN: {filepath}")
|
||||
orphan_count += 1
|
||||
orphan_bytes += file_size
|
||||
else:
|
||||
log(f"JUNK: {filepath}")
|
||||
junk_count += 1
|
||||
junk_bytes += file_size
|
||||
|
||||
total_del_bytes = orphan_bytes + junk_bytes
|
||||
total_removed = orphan_count + junk_count
|
||||
|
||||
# ── Safety gate 7 — deletion size threshold ───────────────────────────────
|
||||
if total_del_bytes > max_del_bytes:
|
||||
total_human = format_bytes(total_del_bytes)
|
||||
if not args.i_know:
|
||||
print()
|
||||
error(f"Deletion would exceed {max_del_gb}GB — {total_human} would be deleted")
|
||||
error("Review ORPHAN lines above carefully before proceeding")
|
||||
error("Rerun with: --i-know-what-im-doing")
|
||||
error("To also bypass age check: add --skip-strike-list")
|
||||
_notify(f"{arr_label} cleanup halted on {hostname} — {total_human} requires --i-know-what-im-doing")
|
||||
sys.exit(1)
|
||||
else:
|
||||
warn(f"OVERRIDE — deletion is {total_human} — proceeding with --i-know-what-im-doing")
|
||||
|
||||
# ── Execute deletions ──────────────────────────────────────────────────────
|
||||
if not args.dry_run:
|
||||
for filepath in all_files:
|
||||
if filepath in tracked:
|
||||
continue
|
||||
if is_protected(filepath, protected):
|
||||
continue
|
||||
try:
|
||||
st = Path(filepath).stat()
|
||||
file_age = now - st.st_mtime
|
||||
except OSError:
|
||||
continue
|
||||
|
||||
if is_media_file(filepath, extensions):
|
||||
if file_age < age_threshold and not args.skip_strikes:
|
||||
continue
|
||||
|
||||
try:
|
||||
Path(filepath).unlink()
|
||||
except OSError as e:
|
||||
error(f"Failed to delete: {filepath} — {e}")
|
||||
|
||||
log("Cleaning up empty folders...")
|
||||
for root in scan_roots:
|
||||
if Path(root).is_dir():
|
||||
for d in sorted(Path(root).rglob("*"), key=lambda p: len(p.parts), reverse=True):
|
||||
if d.is_dir():
|
||||
try:
|
||||
d.rmdir()
|
||||
except OSError:
|
||||
pass
|
||||
log("Empty folders removed")
|
||||
|
||||
elapsed = int(time.time() - start)
|
||||
|
||||
# ── Summary ────────────────────────────────────────────────────────────────
|
||||
orphan_human = format_bytes(orphan_bytes)
|
||||
junk_human = format_bytes(junk_bytes)
|
||||
|
||||
print()
|
||||
print(f"━━━━━ 📋 {arr_label.upper()} CLEANUP SUMMARY ━━━━━")
|
||||
print(f"🖥️ Identity: {my_id} ({server_name})")
|
||||
print(f"🔄 Tracked: {tracked_count} files ({parent_count} {profile['parent_label']})")
|
||||
print(f"🛡️ Protected: {protected_count} files (cover art, metadata)")
|
||||
print(f"🗑️ Orphans: {orphan_count} files ({orphan_human})")
|
||||
print(f"🗑️ Junk: {junk_count} files ({junk_human})")
|
||||
print(f"⏭️ Recent skipped: {recent_count} files (under {orphan_age} days)")
|
||||
print(f"⏱️ Duration: {format_duration(elapsed)}")
|
||||
print()
|
||||
|
||||
if args.dry_run:
|
||||
warn("DRY RUN — no files deleted")
|
||||
elif total_removed == 0:
|
||||
success("Clean — nothing to remove")
|
||||
else:
|
||||
warn(f"🏁 Removed {total_removed} files (orphans: {orphan_human} junk: {junk_human})")
|
||||
_notify(
|
||||
f"{arr_label} cleanup on {hostname} — removed {total_removed} files "
|
||||
f"(orphans: {orphan_human} junk: {junk_human})"
|
||||
)
|
||||
notify_emby_scan(emby_url, emby_api_key, my_id)
|
||||
|
||||
print("━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━")
|
||||
|
||||
# ── Write stats for coffee report ─────────────────────────────────────────
|
||||
if not args.dry_run and stats_file:
|
||||
today = datetime.date.today().strftime("%Y-%m-%d")
|
||||
line = f"{today}|{ARR}|{orphan_count}|{orphan_bytes}|{junk_count}|{junk_bytes}|{recent_count}|{tracked_count}\n"
|
||||
try:
|
||||
with open(stats_file, "a") as f:
|
||||
f.write(line)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,116 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ================================= Arr Cleanup Launcher =======================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Thin bash launcher for arr_cleanup.py. Handles everything bash is uniquely
|
||||
# suited for: sourcing shell config, detect_hosts(), exporting env vars.
|
||||
# All logic lives in Python.
|
||||
#
|
||||
# USAGE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# arr_cleanup.sh --arr lidarr [--dry-run] [--log] [--status]
|
||||
# arr_cleanup.sh --arr radarr [--i-know-what-im-doing] [--skip-strike-list]
|
||||
# arr_cleanup.sh --arr sonarr
|
||||
#
|
||||
# ADDING A NEW ARR
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# 1. Add HOST*_<ARR>_URL, API_KEY, MEDIA_ROOT, PATH_MAP to host*.conf
|
||||
# 2. Add <ARR>_ORPHAN_AGE, MAX_DELETE_GB, EXTENSIONS, etc. to master.conf
|
||||
# 3. Add a profile entry to ARR_PROFILES in arr_cleanup.py (6 values)
|
||||
# 4. Add an export block for the new arr below (copy Sonarr block, change prefix)
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
echo "ERROR: Must be run as root" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
if ! command -v python3 >/dev/null 2>&1; then
|
||||
echo "ERROR: python3 not found — required for arr_cleanup" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
# ── Helper: serialize host-specific associative array to JSON ──────────────────
|
||||
# Reads ${MY_ID}_${ARR_UPPER}_PATH_MAP and emits {"container_path":"host_path",...}
|
||||
# Paths with double-quotes in names are not supported (not a real-world constraint).
|
||||
_path_map_json() {
|
||||
local arr_upper="$1"
|
||||
local map_var="${MY_ID}_${arr_upper}_PATH_MAP"
|
||||
local json="{"
|
||||
local sep="" keys k v
|
||||
|
||||
if ! declare -p "$map_var" 2>/dev/null | grep -q "declare -A"; then
|
||||
echo "{}"
|
||||
return
|
||||
fi
|
||||
|
||||
eval "keys=(\"\${!${map_var}[@]}\")"
|
||||
for k in "${keys[@]}"; do
|
||||
eval "v=\"\${${map_var}[\$k]}\""
|
||||
json+="${sep}\"${k}\":\"${v}\""
|
||||
sep=","
|
||||
done
|
||||
json+="}"
|
||||
echo "$json"
|
||||
}
|
||||
|
||||
# ── Host identity ──────────────────────────────────────────────────────────────
|
||||
export MY_ID LOCAL_SERVER_NAME
|
||||
export ARR_HOSTNAME
|
||||
ARR_HOSTNAME=$(hostname)
|
||||
|
||||
# ── Notifications ──────────────────────────────────────────────────────────────
|
||||
export NOTIFY_UNRAID
|
||||
export EMBY_URL EMBY_API_KEY
|
||||
export MY_DISCORD_WEBHOOK
|
||||
|
||||
# ── Shared ─────────────────────────────────────────────────────────────────────
|
||||
export ARR_CLEANUP_STATS
|
||||
export LOCK_DIR LOCK_WAIT_TIMEOUT
|
||||
|
||||
# ── Lidarr ─────────────────────────────────────────────────────────────────────
|
||||
export LIDARR_URL LIDARR_API_KEY
|
||||
export LIDARR_MEDIA_ROOT="${LIDARR_MUSIC_ROOT:-}"
|
||||
export LIDARR_EXTENSIONS="${LIDARR_EXTENSIONS[*]:-}"
|
||||
export LIDARR_PROTECTED_PATTERNS="${LIDARR_PROTECTED_PATTERNS[*]:-}"
|
||||
export LIDARR_ORPHAN_AGE LIDARR_MAX_DELETE_GB
|
||||
export LIDARR_MIN_TRACKED_PCT LIDARR_TRACKED_COUNT_FILE
|
||||
export LIDARR_VERSION_MAJOR LIDARR_IMPORT_SCAN_TIMEOUT LIDARR_LOCK_WARN_AGE
|
||||
export LIDARR_PATH_MAP_JSON
|
||||
LIDARR_PATH_MAP_JSON=$(_path_map_json "LIDARR")
|
||||
|
||||
# ── Radarr ─────────────────────────────────────────────────────────────────────
|
||||
export RADARR_URL RADARR_API_KEY
|
||||
export RADARR_MEDIA_ROOT="${RADARR_MOVIES_ROOT:-}"
|
||||
export RADARR_EXTENSIONS="${RADARR_EXTENSIONS[*]:-}"
|
||||
export RADARR_PROTECTED_PATTERNS="${RADARR_PROTECTED_PATTERNS[*]:-}"
|
||||
export RADARR_ORPHAN_AGE RADARR_MAX_DELETE_GB
|
||||
export RADARR_MIN_TRACKED_PCT RADARR_TRACKED_COUNT_FILE
|
||||
export RADARR_VERSION_MAJOR RADARR_IMPORT_SCAN_TIMEOUT RADARR_LOCK_WARN_AGE
|
||||
export RADARR_PATH_MAP_JSON
|
||||
RADARR_PATH_MAP_JSON=$(_path_map_json "RADARR")
|
||||
|
||||
# ── Sonarr ─────────────────────────────────────────────────────────────────────
|
||||
export SONARR_URL SONARR_API_KEY
|
||||
export SONARR_MEDIA_ROOT="${SONARR_TV_ROOT:-}"
|
||||
export SONARR_EXTENSIONS="${SONARR_EXTENSIONS[*]:-}"
|
||||
export SONARR_PROTECTED_PATTERNS="${SONARR_PROTECTED_PATTERNS[*]:-}"
|
||||
export SONARR_ORPHAN_AGE SONARR_MAX_DELETE_GB
|
||||
export SONARR_MIN_TRACKED_PCT SONARR_TRACKED_COUNT_FILE
|
||||
export SONARR_VERSION_MAJOR SONARR_IMPORT_SCAN_TIMEOUT SONARR_LOCK_WARN_AGE
|
||||
export SONARR_PATH_MAP_JSON
|
||||
SONARR_PATH_MAP_JSON=$(_path_map_json "SONARR")
|
||||
|
||||
exec python3 "$SCRIPT_DIR/arr_cleanup.py" "$@"
|
||||
@@ -1,517 +0,0 @@
|
||||
#!/bin/bash
|
||||
# ==============================================================================================
|
||||
# ========================= Continuous Scripts Status ==========================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Live status dashboard for all continuously running scripts in the ecosystem.
|
||||
# Run manually at any time — no schedule, no cron.
|
||||
#
|
||||
# For each script shows: running state, PID, uptime, approximate cycle count,
|
||||
# active strikes, skip list, recent restart history, and a live health snapshot.
|
||||
#
|
||||
# system_watchdog — rootfs, RAM, ZFS ARC, load, zombie count, CPU temp
|
||||
# docker_watchdog — running/stopped/unhealthy containers, required containers,
|
||||
# memory-monitored containers, recent restart history
|
||||
# failover — current state, tier status, remote Tailscale visibility
|
||||
#
|
||||
# If a script is mid-cycle, state files are read as-is — reflects last completed cycle.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Host-Aware Output
|
||||
# detect_hosts() sets MY_ID and aliases all HOST*_WATCHDOG_* arrays.
|
||||
# Required containers and tier delays are shown for the correct host.
|
||||
#
|
||||
# Read-Only
|
||||
# Reads state files and docker inspect output only — makes no changes to any
|
||||
# running script, container, or state file.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# continuous_scripts_status.sh
|
||||
# Show the full dashboard for all continuous scripts.
|
||||
#
|
||||
# continuous_scripts_status.sh --log
|
||||
# Verbose output with additional detail per script section.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
# Dashboard script — output is the point
|
||||
SILENT_MODE=false
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
DOCKER_TIMEOUT=15
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
# detect_hosts() sets MY_ID and aliases all HOST*_WATCHDOG_* arrays
|
||||
detect_hosts
|
||||
|
||||
# ==============================================================================================
|
||||
# ── HELPER FUNCTIONS ──────────────────────────────────────────────────────────────────────────
|
||||
# ==============================================================================================
|
||||
|
||||
get_lock_pid() {
|
||||
local script_name="$1"
|
||||
local lockfile="$LOCK_DIR/${script_name}.lock"
|
||||
if [[ -f "$lockfile" ]]; then
|
||||
local content
|
||||
content=$(cat "$lockfile" 2>/dev/null)
|
||||
echo "${content%%:*}"
|
||||
fi
|
||||
}
|
||||
|
||||
get_lock_name() {
|
||||
local script_name="$1"
|
||||
local lockfile="$LOCK_DIR/${script_name}.lock"
|
||||
if [[ -f "$lockfile" ]]; then
|
||||
local content
|
||||
content=$(cat "$lockfile" 2>/dev/null)
|
||||
echo "${content##*:}"
|
||||
fi
|
||||
}
|
||||
|
||||
is_script_running() {
|
||||
local script_name="$1"
|
||||
local pid locked_name
|
||||
pid=$(get_lock_pid "$script_name")
|
||||
locked_name=$(get_lock_name "$script_name")
|
||||
[[ -n "$pid" ]] && kill -0 "$pid" 2>/dev/null && [[ "$locked_name" == "$script_name" ]]
|
||||
}
|
||||
|
||||
get_lock_age() {
|
||||
local script_name="$1"
|
||||
local lockfile="$LOCK_DIR/${script_name}.lock"
|
||||
if [[ -f "$lockfile" ]]; then
|
||||
local mtime now
|
||||
mtime=$(stat -c %Y "$lockfile" 2>/dev/null || echo 0)
|
||||
now=$(date +%s)
|
||||
echo $(( now - mtime ))
|
||||
else
|
||||
echo 0
|
||||
fi
|
||||
}
|
||||
|
||||
# Human readable uptime — days/hours/mins
|
||||
format_uptime() {
|
||||
local seconds=$1
|
||||
local days=$(( seconds / 86400 ))
|
||||
local hours=$(( (seconds % 86400) / 3600 ))
|
||||
local mins=$(( (seconds % 3600) / 60 ))
|
||||
if (( days > 0 )); then
|
||||
echo "${days}d ${hours}h ${mins}m"
|
||||
elif (( hours > 0 )); then
|
||||
echo "${hours}h ${mins}m"
|
||||
else
|
||||
echo "${mins}m"
|
||||
fi
|
||||
}
|
||||
|
||||
divider() { printf '%.0s─' {1..57}; echo; }
|
||||
section() { echo ""; echo " $1"; divider; }
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Header ━━━
|
||||
# ==============================================================================================
|
||||
clear
|
||||
echo ""
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo " 🛡️ WATCHDOG STATUS — $(date '+%A, %B %-d at %-I:%M%p')"
|
||||
echo " $ICON_HOST $MY_ID — $LOCAL_SERVER_NAME"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ System Watchdog ━━━
|
||||
# ==============================================================================================
|
||||
section "⚙️ SYSTEM WATCHDOG"
|
||||
|
||||
SYS_PID=$(get_lock_pid "system_watchdog")
|
||||
SYS_RUNNING=false
|
||||
|
||||
if is_script_running "system_watchdog"; then
|
||||
SYS_RUNNING=true
|
||||
SYS_AGE=$(get_lock_age "system_watchdog")
|
||||
SYS_UPTIME=$(format_uptime "$SYS_AGE")
|
||||
SYS_CYCLE=$(( SYS_AGE / SYSTEM_WATCHDOG_INTERVAL ))
|
||||
echo " ✅ Running │ PID: $SYS_PID │ Uptime: $SYS_UPTIME │ ~Cycle: $SYS_CYCLE"
|
||||
echo " ⏱️ Interval: ${SYSTEM_WATCHDOG_INTERVAL}s │ Heartbeat every: ${SYSTEM_WATCHDOG_HEARTBEAT_HOURS}hr"
|
||||
else
|
||||
echo " ❌ NOT RUNNING — system_watchdog.sh is not active"
|
||||
echo " Start via: bash Orchestrators/array_started.sh"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
|
||||
# System strikes
|
||||
if [[ -f "$SYS_WATCHDOG_STATE_FILE" ]]; then
|
||||
ACTIVE_STRIKES=$(grep -v ":0$" "$SYS_WATCHDOG_STATE_FILE" 2>/dev/null | grep -v "^$")
|
||||
if [[ -n "$ACTIVE_STRIKES" ]]; then
|
||||
echo " ⚠️ Active strikes:"
|
||||
while IFS=: read -r key count; do
|
||||
[[ -z "$key" ]] && continue
|
||||
echo " → $key: $count/$SYS_WATCHDOG_STRIKE_LIMIT"
|
||||
done <<< "$ACTIVE_STRIKES"
|
||||
else
|
||||
echo " ✅ Strikes: none"
|
||||
fi
|
||||
else
|
||||
echo " ℹ️ Strike state file not found (watchdog may not have run yet)"
|
||||
fi
|
||||
|
||||
# Reboot log
|
||||
if [[ -f "$SYS_WATCHDOG_REBOOT_LOG" ]]; then
|
||||
TOTAL_REBOOTS=$(grep -c "." "$SYS_WATCHDOG_REBOOT_LOG" 2>/dev/null || echo 0)
|
||||
TOTAL_REBOOTS="${TOTAL_REBOOTS//[^0-9]/}"
|
||||
TOTAL_REBOOTS="${TOTAL_REBOOTS:-0}"
|
||||
WEEK_CUTOFF=$(date -d "7 days ago" '+%Y-%m-%d %H:%M:%S')
|
||||
WEEK_REBOOTS=$(awk -v cutoff="$WEEK_CUTOFF" '$0 >= cutoff' \
|
||||
"$SYS_WATCHDOG_REBOOT_LOG" 2>/dev/null | wc -l)
|
||||
echo " 🔄 Watchdog reboots: $WEEK_REBOOTS this week / $TOTAL_REBOOTS total"
|
||||
fi
|
||||
|
||||
# Container skip list
|
||||
if [[ -f "$SYS_WATCHDOG_FAILED_FILE" ]] && [[ -s "$SYS_WATCHDOG_FAILED_FILE" ]]; then
|
||||
SKIP_COUNT=$(wc -l < "$SYS_WATCHDOG_FAILED_FILE")
|
||||
echo ""
|
||||
echo " ⛔ Skip list ($SKIP_COUNT — manual intervention needed):"
|
||||
while IFS= read -r container; do
|
||||
[[ -z "$container" ]] && continue
|
||||
echo " → $container"
|
||||
done < "$SYS_WATCHDOG_FAILED_FILE"
|
||||
else
|
||||
echo " ✅ Skip list: empty"
|
||||
fi
|
||||
|
||||
# Live system health snapshot
|
||||
echo ""
|
||||
echo " 📊 Current system state:"
|
||||
|
||||
ROOTFS_PCT=$(df / --output=pcent 2>/dev/null | tail -1 | tr -d ' %')
|
||||
[[ "${ROOTFS_PCT:-0}" -ge "${SYS_WATCHDOG_ROOTFS_PCT:-95}" ]] && \
|
||||
ROOTFS_ICON="⚠️ " || ROOTFS_ICON="✅"
|
||||
echo " ${ROOTFS_ICON} rootfs: ${ROOTFS_PCT}% (threshold: ${SYS_WATCHDOG_ROOTFS_PCT}%)"
|
||||
|
||||
MEM_AVAIL_KB=$(awk '/MemAvailable/ {print $2}' /proc/meminfo)
|
||||
MEM_FREE_GB=$(awk "BEGIN {printf \"%.1f\", $MEM_AVAIL_KB / 1048576}")
|
||||
MEM_TOTAL_GB=$(awk '/MemTotal/ {printf "%.0f", $2/1048576}' /proc/meminfo)
|
||||
[[ $(printf "%.0f" "$MEM_FREE_GB") -lt "${SYS_WATCHDOG_MEM_GB:-4}" ]] && \
|
||||
MEM_ICON="⚠️ " || MEM_ICON="✅"
|
||||
echo " ${MEM_ICON} RAM: ${MEM_FREE_GB}GB free / ${MEM_TOTAL_GB}GB total (threshold: ${SYS_WATCHDOG_MEM_GB}GB free)"
|
||||
|
||||
if [[ -f /proc/spl/kstat/zfs/arcstats ]]; then
|
||||
ARC_SIZE=$(awk '/^size / {print $3}' /proc/spl/kstat/zfs/arcstats)
|
||||
ARC_MAX=$(awk '/^c_max / {print $3}' /proc/spl/kstat/zfs/arcstats)
|
||||
ARC_PCT=$(( ARC_SIZE * 100 / ARC_MAX ))
|
||||
ARC_GB=$(awk "BEGIN {printf \"%.1f\", $ARC_SIZE / 1073741824}")
|
||||
[[ "$ARC_PCT" -ge "${SYS_WATCHDOG_ARC_PINNED_PCT:-98}" ]] && \
|
||||
ARC_ICON="⚠️ " || ARC_ICON="✅"
|
||||
echo " ${ARC_ICON} ZFS ARC: ${ARC_GB}GB (${ARC_PCT}% of max, threshold: ${SYS_WATCHDOG_ARC_PINNED_PCT}%)"
|
||||
fi
|
||||
|
||||
LOAD=$(awk '{print $1}' /proc/loadavg)
|
||||
CORES=$(nproc)
|
||||
LOAD_THRESH=$(( CORES * ${SYS_WATCHDOG_LOAD_MULTIPLIER:-3} ))
|
||||
LOAD_INT=$(printf "%.0f" "$LOAD")
|
||||
[[ "$LOAD_INT" -ge "$LOAD_THRESH" ]] && LOAD_ICON="⚠️ " || LOAD_ICON="✅"
|
||||
echo " ${LOAD_ICON} Load avg: $LOAD (threshold: ${LOAD_THRESH} = ${SYS_WATCHDOG_LOAD_MULTIPLIER}x ${CORES} cores)"
|
||||
|
||||
ZOMBIE_COUNT=$(ps aux | awk '{print $8}' | grep -c "^Z$" 2>/dev/null || echo 0)
|
||||
ZOMBIE_COUNT="${ZOMBIE_COUNT//[^0-9]/}"
|
||||
ZOMBIE_COUNT="${ZOMBIE_COUNT:-0}"
|
||||
[[ "$ZOMBIE_COUNT" -ge "${SYS_WATCHDOG_ZOMBIE_LIMIT:-50}" ]] && \
|
||||
ZOMBIE_ICON="⚠️ " || ZOMBIE_ICON="✅"
|
||||
echo " ${ZOMBIE_ICON} Zombies: $ZOMBIE_COUNT (threshold: ${SYS_WATCHDOG_ZOMBIE_LIMIT})"
|
||||
|
||||
if command -v sensors >/dev/null 2>&1; then
|
||||
CPU_TEMP=$(sensors 2>/dev/null | \
|
||||
grep -i "Package id 0\|Tctl\|CPU Temp" | \
|
||||
awk '{print $NF}' | tr -d '+°C' | head -1)
|
||||
if [[ -n "$CPU_TEMP" ]]; then
|
||||
CPU_TEMP_INT=$(printf "%.0f" "$CPU_TEMP")
|
||||
[[ "$CPU_TEMP_INT" -ge "${SYS_WATCHDOG_CPU_TEMP_MAX:-95}" ]] && \
|
||||
TEMP_ICON="⚠️ " || TEMP_ICON="✅"
|
||||
echo " ${TEMP_ICON} CPU temp: ${CPU_TEMP_INT}°C (threshold: ${SYS_WATCHDOG_CPU_TEMP_MAX}°C)"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Docker Watchdog ━━━
|
||||
# ==============================================================================================
|
||||
section "🐳 DOCKER WATCHDOG"
|
||||
|
||||
DOCKER_PID=$(get_lock_pid "docker_watchdog")
|
||||
DOCKER_RUNNING=false
|
||||
|
||||
if is_script_running "docker_watchdog"; then
|
||||
DOCKER_RUNNING=true
|
||||
DOCKER_AGE=$(get_lock_age "docker_watchdog")
|
||||
DOCKER_UPTIME=$(format_uptime "$DOCKER_AGE")
|
||||
DOCKER_CYCLE=$(( DOCKER_AGE / DOCKER_WATCHDOG_INTERVAL ))
|
||||
echo " ✅ Running │ PID: $DOCKER_PID │ Uptime: $DOCKER_UPTIME │ ~Cycle: $DOCKER_CYCLE"
|
||||
echo " ⏱️ Interval: ${DOCKER_WATCHDOG_INTERVAL}s │ Heartbeat every: ${DOCKER_WATCHDOG_HEARTBEAT_HOURS}hr"
|
||||
else
|
||||
echo " ❌ NOT RUNNING — docker_watchdog.sh is not active"
|
||||
echo " Start via: bash Orchestrators/array_started.sh"
|
||||
fi
|
||||
|
||||
echo ""
|
||||
|
||||
# Container strikes
|
||||
if [[ -f "$WATCHDOG_STATE_FILE" ]]; then
|
||||
ACTIVE_CONTAINER_STRIKES=$(grep -v ":0$" "$WATCHDOG_STATE_FILE" 2>/dev/null | grep -v "^$")
|
||||
if [[ -n "$ACTIVE_CONTAINER_STRIKES" ]]; then
|
||||
echo " ⚠️ Active container strikes:"
|
||||
while IFS=: read -r key count; do
|
||||
[[ -z "$key" ]] && continue
|
||||
echo " → $key: $count"
|
||||
done <<< "$ACTIVE_CONTAINER_STRIKES"
|
||||
else
|
||||
echo " ✅ Container strikes: none"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Container restart history
|
||||
if [[ -f "$WATCHDOG_CONTAINER_RESTART_LOG" ]]; then
|
||||
WEEK_CUTOFF=$(date -d "7 days ago" '+%Y-%m-%d %H:%M:%S')
|
||||
WEEK_RESTARTS=$(awk -F'|' -v cutoff="$WEEK_CUTOFF" \
|
||||
'$2 >= cutoff' "$WATCHDOG_CONTAINER_RESTART_LOG" 2>/dev/null | wc -l)
|
||||
if [[ "${WEEK_RESTARTS:-0}" -gt 0 ]]; then
|
||||
echo ""
|
||||
echo " 🔄 Container restarts this week: $WEEK_RESTARTS"
|
||||
awk -F'|' -v cutoff="$WEEK_CUTOFF" \
|
||||
'$2 >= cutoff {print $1}' "$WATCHDOG_CONTAINER_RESTART_LOG" 2>/dev/null | \
|
||||
sort | uniq -c | sort -rn | head -5 | \
|
||||
while read -r count name; do
|
||||
echo " → $name: $count restart(s)"
|
||||
done
|
||||
else
|
||||
echo " ✅ Container restarts this week: none"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Container overview
|
||||
echo ""
|
||||
echo " 📦 Container overview:"
|
||||
|
||||
if command -v docker >/dev/null 2>&1; then
|
||||
RUNNING=$(timeout "$DOCKER_TIMEOUT" docker ps -q 2>/dev/null | wc -l)
|
||||
TOTAL=$(timeout "$DOCKER_TIMEOUT" docker ps -aq 2>/dev/null | wc -l)
|
||||
UNHEALTHY=$(timeout "$DOCKER_TIMEOUT" docker ps \
|
||||
--filter health=unhealthy -q 2>/dev/null | wc -l)
|
||||
|
||||
# Stopped containers — bucket into clean vs unexpected, skip SCAN_IGNORE entirely
|
||||
CLEAN_STOPPED=()
|
||||
UNEXPECTED_STOPPED=()
|
||||
while IFS= read -r name; do
|
||||
[[ -z "$name" ]] && continue
|
||||
SKIP=false
|
||||
for ignore in "${WATCHDOG_SCAN_IGNORE[@]}"; do
|
||||
[[ "$name" == "$ignore" ]] && SKIP=true && break
|
||||
done
|
||||
[[ "$SKIP" == true ]] && continue
|
||||
exit_code=$(docker inspect --format '{{.State.ExitCode}}' "$name" 2>/dev/null)
|
||||
if [[ "$exit_code" == "0" || "$exit_code" == "143" ]]; then
|
||||
CLEAN_STOPPED+=("$name")
|
||||
else
|
||||
UNEXPECTED_STOPPED+=("$name")
|
||||
fi
|
||||
done < <(timeout "$DOCKER_TIMEOUT" docker ps -af "status=exited" \
|
||||
--format "{{.Names}}" 2>/dev/null)
|
||||
|
||||
echo " Running: $RUNNING / $TOTAL total"
|
||||
[[ "$UNHEALTHY" -gt 0 ]] && echo " ⚠️ Unhealthy: $UNHEALTHY"
|
||||
|
||||
if [[ "${#UNEXPECTED_STOPPED[@]}" -gt 0 ]]; then
|
||||
echo " ⚠️ Stopped (unexpected):"
|
||||
for name in "${UNEXPECTED_STOPPED[@]}"; do
|
||||
echo " → $name"
|
||||
done
|
||||
fi
|
||||
|
||||
if [[ "${#CLEAN_STOPPED[@]}" -gt 0 ]]; then
|
||||
echo " ⏸️ Stopped (clean):"
|
||||
for name in "${CLEAN_STOPPED[@]}"; do
|
||||
echo " → $name"
|
||||
done
|
||||
fi
|
||||
|
||||
if [[ "${#UNEXPECTED_STOPPED[@]}" -eq 0 && "${#CLEAN_STOPPED[@]}" -eq 0 ]]; then
|
||||
echo " ✅ All containers running"
|
||||
fi
|
||||
|
||||
# Required containers — aliased by detect_hosts() → WATCHDOG_REQUIRED_CONTAINERS
|
||||
REQUIRED_ISSUES=0
|
||||
if [[ ${#WATCHDOG_REQUIRED_CONTAINERS[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
echo " 🔐 Required containers:"
|
||||
for container in "${WATCHDOG_REQUIRED_CONTAINERS[@]}"; do
|
||||
[[ -z "$container" ]] && continue
|
||||
STATUS=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Running}}' "$container" 2>/dev/null || echo "not found")
|
||||
if [[ "$STATUS" == "true" ]]; then
|
||||
echo " ✅ $container"
|
||||
else
|
||||
echo " ❌ $container — $STATUS"
|
||||
(( REQUIRED_ISSUES++ ))
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
# Memory-monitored containers — aliased by detect_hosts() → WATCHDOG_CONTAINERS
|
||||
if [[ ${#WATCHDOG_CONTAINERS[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
echo " 📊 Monitored containers (memory):"
|
||||
for container in "${!WATCHDOG_CONTAINERS[@]}"; do
|
||||
LIMIT_MB="${WATCHDOG_CONTAINERS[$container]}"
|
||||
LIMIT_GB=$(awk "BEGIN {printf \"%.0f\", $LIMIT_MB / 1024}")
|
||||
USAGE=$(timeout "$DOCKER_TIMEOUT" docker stats --no-stream \
|
||||
--format "{{.MemUsage}}" "$container" 2>/dev/null | awk '{print $1}')
|
||||
STATUS=$(timeout "$DOCKER_TIMEOUT" docker inspect -f \
|
||||
'{{.State.Running}}' "$container" 2>/dev/null || echo "not found")
|
||||
if [[ "$STATUS" == "true" ]]; then
|
||||
echo " ✅ $container: ${USAGE:-?} (limit: ${LIMIT_GB}GB)"
|
||||
else
|
||||
echo " ❌ $container: not running (limit: ${LIMIT_GB}GB)"
|
||||
fi
|
||||
done
|
||||
fi
|
||||
else
|
||||
echo " Docker not available"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Failover ━━━
|
||||
# ==============================================================================================
|
||||
section "🔀 FALLBACK"
|
||||
|
||||
FALLBACK_PID=$(get_lock_pid "fallback")
|
||||
FALLBACK_RUNNING=false
|
||||
|
||||
if is_script_running "fallback"; then
|
||||
FALLBACK_RUNNING=true
|
||||
FALLBACK_AGE=$(get_lock_age "fallback")
|
||||
FALLBACK_UPTIME=$(format_uptime "$FALLBACK_AGE")
|
||||
echo " ✅ Running │ PID: $FALLBACK_PID │ Uptime: $FALLBACK_UPTIME"
|
||||
else
|
||||
if [[ "${FALLBACK_ENABLED:-true}" == false ]]; then
|
||||
echo " ⏸️ Disabled — FALLBACK_ENABLED=false in master.conf"
|
||||
else
|
||||
echo " ❌ NOT RUNNING — fallback.sh is not active"
|
||||
echo " Start via: bash Orchestrators/array_started.sh"
|
||||
fi
|
||||
fi
|
||||
|
||||
echo ""
|
||||
|
||||
# Fallback state
|
||||
FALLBACK_STATE="UNKNOWN"
|
||||
FALLBACK_STATE_SECONDS=0
|
||||
|
||||
if [[ -f "$FALLBACK_STATE_FILE" ]]; then
|
||||
FALLBACK_STATE=$(grep "^state=" "$FALLBACK_STATE_FILE" 2>/dev/null | cut -d= -f2)
|
||||
FALLBACK_LAST_EPOCH=$(grep "^fallback_start=" "$FALLBACK_STATE_FILE" \
|
||||
2>/dev/null | cut -d= -f2)
|
||||
if [[ -n "$FALLBACK_LAST_EPOCH" && "$FALLBACK_LAST_EPOCH" -gt 0 ]]; then
|
||||
FALLBACK_STATE_SECONDS=$(( $(date +%s) - FALLBACK_LAST_EPOCH ))
|
||||
fi
|
||||
fi
|
||||
|
||||
STATE_DURATION=$(format_uptime "${FALLBACK_STATE_SECONDS:-0}")
|
||||
|
||||
# Tier delays via REMOTE_ID — same logic as fallback.sh
|
||||
REMOTE_TIER2_VAR="${REMOTE_ID}_TIER2_DELAY"
|
||||
REMOTE_TIER3_VAR="${REMOTE_ID}_TIER3_DELAY"
|
||||
REMOTE_TIER4_VAR="${REMOTE_ID}_TIER4_DELAY"
|
||||
TIER2_DELAY="${!REMOTE_TIER2_VAR:-240}"
|
||||
TIER3_DELAY="${!REMOTE_TIER3_VAR:-720}"
|
||||
TIER4_DELAY="${!REMOTE_TIER4_VAR:-1440}"
|
||||
|
||||
case "$FALLBACK_STATE" in
|
||||
NORMAL)
|
||||
echo " ✅ State: NORMAL"
|
||||
;;
|
||||
FALLBACK)
|
||||
echo " ⚠️ State: FALLBACK — $REMOTE_SERVER_NAME is down"
|
||||
echo " ⏱️ Duration: $STATE_DURATION"
|
||||
FALLBACK_MINS=$(( FALLBACK_STATE_SECONDS / 60 ))
|
||||
echo ""
|
||||
echo " 🔄 Tier status:"
|
||||
echo " Tier 1 (immediate): ✅ active"
|
||||
if (( FALLBACK_MINS >= TIER2_DELAY )); then
|
||||
echo " Tier 2 (${TIER2_DELAY}min): ✅ active"
|
||||
else
|
||||
REMAINING=$(( TIER2_DELAY - FALLBACK_MINS ))
|
||||
echo " Tier 2 (${TIER2_DELAY}min): ⏳ in ${REMAINING}min"
|
||||
fi
|
||||
if (( FALLBACK_MINS >= TIER3_DELAY )); then
|
||||
echo " Tier 3 (${TIER3_DELAY}min): ✅ active"
|
||||
else
|
||||
REMAINING=$(( TIER3_DELAY - FALLBACK_MINS ))
|
||||
echo " Tier 3 (${TIER3_DELAY}min): ⏳ in ${REMAINING}min"
|
||||
fi
|
||||
if (( FALLBACK_MINS >= TIER4_DELAY )); then
|
||||
echo " Tier 4 (${TIER4_DELAY}min): ✅ active"
|
||||
else
|
||||
REMAINING=$(( TIER4_DELAY - FALLBACK_MINS ))
|
||||
echo " Tier 4 (${TIER4_DELAY}min): ⏳ in ${REMAINING}min"
|
||||
fi
|
||||
;;
|
||||
NO_INTERNET)
|
||||
echo " ❌ State: NO_INTERNET — DDNS stopped"
|
||||
echo " ⏱️ Down for: $STATE_DURATION"
|
||||
;;
|
||||
DARK)
|
||||
echo " ❌ State: DARK — $REMOTE_SERVER_NAME down AND no internet"
|
||||
echo " ⏱️ Duration: $STATE_DURATION"
|
||||
;;
|
||||
*)
|
||||
echo " ❓ State: ${FALLBACK_STATE:-unknown}"
|
||||
;;
|
||||
esac
|
||||
|
||||
|
||||
echo " 📡 Check interval: ${FALLBACK_CHECK_INTERVAL}s │ Handback strikes: ${FALLBACK_HANDBACK_STRIKES}"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Footer ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
ISSUES=0
|
||||
[[ "$SYS_RUNNING" == false ]] && (( ISSUES++ ))
|
||||
[[ "$DOCKER_RUNNING" == false ]] && (( ISSUES++ ))
|
||||
[[ "$FALLBACK_RUNNING" == false && "${FALLBACK_ENABLED:-true}" != false ]] && (( ISSUES++ ))
|
||||
[[ -n "$ACTIVE_STRIKES" ]] && (( ISSUES++ ))
|
||||
[[ -n "$ACTIVE_CONTAINER_STRIKES" ]] && (( ISSUES++ ))
|
||||
[[ "${REQUIRED_ISSUES:-0}" -gt 0 ]] && (( ISSUES++ ))
|
||||
[[ "$FALLBACK_STATE" != "NORMAL" && "$FALLBACK_STATE" != "UNKNOWN" ]] && (( ISSUES++ ))
|
||||
|
||||
if [[ "$ISSUES" -eq 0 ]]; then
|
||||
echo " ✅ $MY_ID — all continuous scripts healthy"
|
||||
else
|
||||
echo " ⚠️ $ISSUES issue(s) detected — review above"
|
||||
fi
|
||||
|
||||
echo " 🕐 Checked: $(date '+%H:%M:%S')"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo ""
|
||||
@@ -7,10 +7,10 @@ Orchestrators contain no business logic — they call other scripts in order, tr
|
||||
pass/fail per job, and produce one clean summary. Configuration lives in `master.conf`.
|
||||
Adding or removing a job never requires touching the orchestrator script itself.
|
||||
|
||||
> **The User Scripts plugin contains only orchestrators.** Every cron entry, every
|
||||
> "At Startup of Array" entry, every scheduled operation runs through an orchestrator.
|
||||
> The individual scripts it calls are never scheduled directly — they run in a defined
|
||||
> order inside a coordinated window, with a unified summary at the end.
|
||||
> **The Varaverk scheduler runs only orchestrators.** Every cron entry, every array
|
||||
> start/stop event, every scheduled operation runs through an orchestrator. The individual
|
||||
> scripts it calls are never scheduled directly — they run in a defined order inside a
|
||||
> coordinated window, with a unified summary at the end.
|
||||
|
||||
---
|
||||
|
||||
@@ -71,7 +71,7 @@ until someone manually opens Lidarr, identifies the problem, blocklists the rele
|
||||
and triggers a new search. This takes minutes to do — but nobody does it at 3am
|
||||
when it usually happens.
|
||||
|
||||
The fix: `arrs_failed_stalled_recovery.sh` runs every 6 hours. It finds all
|
||||
The fix: `arrs_failed_stalled_recovery.sh` runs every 4 hours (via `intermediate_sync_maintenance.sh`). It finds all
|
||||
`importFailed`, `importPending`, `error`, and `stalled` items, blocklists them, removes
|
||||
them from the queue, and triggers a new search — automatically. By morning the failed
|
||||
import has already been replaced by a working one. No manual intervention required.
|
||||
@@ -80,33 +80,36 @@ import has already been replaced by a working one. No manual intervention requir
|
||||
|
||||
### 🔴 Array Start Scripts Running in Wrong Order or Not at All
|
||||
|
||||
Scripts configured in the User Scripts plugin as "At Startup of Array" run in an
|
||||
unpredictable order. The ramdisk setup might run after Emby starts. The syslog
|
||||
filter might run after containers have already created veth interfaces. PHP-FPM
|
||||
tuning might run after the WebGUI has already served its first requests. Each
|
||||
script competes for the same startup slot with no guaranteed order.
|
||||
Startup scripts configured individually ran in an unpredictable order. The ramdisk
|
||||
setup might run after Emby starts. The syslog filter might run after containers have
|
||||
already created veth interfaces. PHP-FPM tuning might run after the WebGUI has already
|
||||
served its first requests. Each script competed for the same startup slot with no
|
||||
guaranteed order.
|
||||
|
||||
The fix: `array_started.sh` is the only "At Startup of Array" entry. It launches
|
||||
every startup script in a defined order, with one-second settle between each, and
|
||||
reports which succeeded and which failed. Order is guaranteed. Nothing starts before
|
||||
its dependency. Everything is visible in a single summary.
|
||||
The fix: `array_started.sh` is the only array-start entry in the Varaverk scheduler.
|
||||
It launches every startup script in a defined order, with one-second settle between
|
||||
each, and reports which succeeded and which failed. Order is guaranteed. Nothing starts
|
||||
before its dependency. Everything is visible in a single summary.
|
||||
|
||||
---
|
||||
|
||||
## ━━━ THE ORCHESTRATOR MODEL ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
The User Scripts plugin contains exactly these entries:
|
||||
The Varaverk scheduler contains exactly these entries:
|
||||
|
||||
```bash
|
||||
# At Startup of Array — single entry for all startup scripts:
|
||||
# Array start event (Varaverk disks_mounted hook → cron: "array_start"):
|
||||
array_started.sh
|
||||
|
||||
# Cron — one entry per maintenance window:
|
||||
# Cron — managed via Varaverk Scheduler:
|
||||
*/7 * * * * transcode_management.sh
|
||||
0 */6 * * * arrs_failed_stalled_recovery.sh
|
||||
*/30 * * * * critical_sync_maintenance.sh ← auth + Emby dirty sync + partnership
|
||||
*/15 * * * * watchdog_orchestrator.sh ← resource → docker → system → stability
|
||||
*/30 * * * * critical_sync_maintenance.sh ← auth + play_state_sync + partnership
|
||||
0 */4 * * * intermediate_sync_maintenance.sh ← arr sync + failed recovery + optional rsync
|
||||
0 1 * * * daily_sync_maintenance.sh
|
||||
0 7 * * 0 sunday_morning_coffee_report.sh
|
||||
30 2 * * 0 weekly_sync_maintenance.sh
|
||||
0 0 15 * * monthly_maintenance.sh ← uptime-gated: ZFS scrub, SMART tests
|
||||
|
||||
# Manual only (not scheduled):
|
||||
fallback_test.sh, emby_database_repair.sh, repair tools
|
||||
@@ -156,7 +159,7 @@ visible. Per-item detail suppressed.
|
||||
phase headers, per-phase completion status, and the final summary are visible. Per-share
|
||||
and per-job detail suppressed.
|
||||
|
||||
**Daemon orchestrators** (`watchdog_orchestrator.sh`, `transcode_management.sh`): silent
|
||||
**High-frequency orchestrators** (`watchdog_orchestrator.sh`, `transcode_management.sh`): silent
|
||||
during clean cycles. Only state transitions, errors, and startup-grace expiry shown
|
||||
without `--log`.
|
||||
|
||||
@@ -166,14 +169,13 @@ without `--log`.
|
||||
|
||||
| Script | What It Orchestrates | Schedule |
|
||||
|--------|---------------------|----------|
|
||||
| `array_started.sh` | All array startup scripts in order | At Startup of Array |
|
||||
| `watchdog_orchestrator.sh` | resource → docker → system → stability watchdogs | Every minute |
|
||||
| `array_started.sh` | All array startup scripts in order | `array_start` event (Varaverk plugin hook) |
|
||||
| `watchdog_orchestrator.sh` | resource → docker → system → api_renew → stability watchdogs | Every 15 minutes |
|
||||
| `transcode_management.sh` | Cleanup then manager — order critical | Every 7 minutes |
|
||||
| `arrs_failed_stalled_recovery.sh` | Failed import + stalled download recovery | Every 6 hours |
|
||||
| `daily_sync_maintenance.sh` | git pull → sync → media maintenance → restarts | 1am daily |
|
||||
| `weekly_sync_maintenance.sh` | Stop → update → clean sync → start → weekly restarts | 2:30am Sunday |
|
||||
| `monthly_maintenance.sh` | Uptime-triggered heavy tasks — ZFS scrub, SMART tests | Daily check, fires when uptime ≥ 30d |
|
||||
| `media_management.sh` | Permissions → cleaners → arr cleanup | Via daily_sync (or manual) |
|
||||
| `intermediate_sync_maintenance.sh` | arr sync + arrs_failed_stalled_recovery + optional rsync | Every 4 hours |
|
||||
|
||||
---
|
||||
|
||||
@@ -181,13 +183,13 @@ without `--log`.
|
||||
## 🚀 array_started.sh
|
||||
## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
Single "At Startup of Array" entry for the entire ecosystem. Launches every startup
|
||||
script in order — each as a background process — and reports which succeeded and which
|
||||
failed. You never need to add individual scripts to the User Scripts startup list.
|
||||
Single array-start entry for the entire ecosystem. Fired by the Varaverk plugin's
|
||||
`disks_mounted` event hook. Launches every startup script in order — each as a
|
||||
background process — and reports which succeeded and which failed.
|
||||
|
||||
```bash
|
||||
# Scheduled: At Startup of Array (User Scripts plugin)
|
||||
# This is the ONLY "At Startup of Array" entry in User Scripts
|
||||
# Triggered by: Plugin/unraid/event/disks_mounted/array_start_jobs
|
||||
# schedule.json entry: "Orchestrators/array_started.sh" → cron: "array_start"
|
||||
```
|
||||
|
||||
---
|
||||
@@ -202,28 +204,32 @@ failed. You never need to add individual scripts to the User Scripts startup lis
|
||||
#
|
||||
ARRAY_START_SCRIPTS=(
|
||||
# ── One-shot scripts — run and exit naturally ─────────────────────────────
|
||||
"unRAID_Essentials/inotify_tuning.sh" # raise inotify BEFORE containers start
|
||||
# containers inherit limits at startup —
|
||||
# if Code-Server starts with low limits
|
||||
# it keeps them until restart
|
||||
"unRAID_Essentials/docker_syslog_filter.sh" # suppress veth noise BEFORE containers create
|
||||
# veth interfaces — otherwise the first boot
|
||||
# always has unfiltered veth spam
|
||||
"unRAID_Essentials/php_fpm_max_children.sh" # WebGUI tuning — before any WebGUI requests
|
||||
"unRAID_Essentials/ramdisk_setup.sh" # create tmpfs + symlink BEFORE Emby starts —
|
||||
"Plugin/unraid/System_Essentials/unraid_api_key_renew.sh" # re-register API key FIRST —
|
||||
# registry is ephemeral,
|
||||
# lost on service restart
|
||||
"System_Essentials/conf_sync.sh" # pull conf from remote BEFORE anything
|
||||
# needs current config
|
||||
"System_Essentials/conf_cache_restore.sh" # restore cached conf if pull failed —
|
||||
# ensures config is always available
|
||||
"Transcodes/ramdisk_setup.sh" # create tmpfs + symlink BEFORE Emby starts —
|
||||
# Emby needs the transcode path to exist
|
||||
"System_Essentials/docker_syslog_filter.sh" # suppress veth noise BEFORE containers create
|
||||
# veth interfaces — otherwise first boot has
|
||||
# unfiltered veth spam
|
||||
"Plugin/unraid/System_Essentials/php_fpm_max_children.sh" # WebGUI tuning —
|
||||
# before any WebGUI requests
|
||||
"System_Essentials/inotify_tuning.sh" # raise inotify BEFORE containers start —
|
||||
# containers inherit limits at startup
|
||||
"Docker_Essentials/docker_network_connect.sh" # ensure networks + connections BEFORE
|
||||
# watchdogs check container states
|
||||
|
||||
# ── Continuous scripts — run until array stops ─────────────────────────────
|
||||
"Watchdogs/stability_watchdog.sh" # last line of defense — reboots when all else fails —
|
||||
# system watchdog writes state file that
|
||||
# docker watchdog reads every cycle
|
||||
"Watchdogs/docker_watchdog.sh" # container health BEFORE failover —
|
||||
# containers must be healthy for failover
|
||||
# to make reliable decisions
|
||||
"Fallback/fallback.sh" # fallback LAST — needs everything else stable
|
||||
"Arrs_Stack/start_webhook_listener.sh" # start webhook listener before arrs POST events
|
||||
"Fallback/fallback.sh" # fallback LAST — needs everything else stable
|
||||
)
|
||||
|
||||
# NOTE: watchdogs (docker_watchdog, system_watchdog, stability_watchdog) are NOT here.
|
||||
# They run via watchdog_orchestrator.sh on cron every 15 minutes — not as daemons.
|
||||
```
|
||||
|
||||
---
|
||||
@@ -264,7 +270,7 @@ ARRAY_START_SCRIPTS=(
|
||||
|
||||
```bash
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Normal — called by User Scripts at array start. Never run manually in production.
|
||||
# Normal — fired by Varaverk disks_mounted event hook. Never run manually in production.
|
||||
# array_started.sh runs once and exits — the continuous scripts it launched
|
||||
# keep running as background processes.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -388,7 +394,7 @@ and Lidarr. Blocklists the bad release, removes it from the queue, and triggers
|
||||
new search — hands-free recovery while you sleep.
|
||||
|
||||
```bash
|
||||
# Scheduled: 0 */6 * * * (every 6 hours)
|
||||
# Called by: intermediate_sync_maintenance.sh (every 4 hours)
|
||||
```
|
||||
|
||||
---
|
||||
@@ -532,7 +538,7 @@ Runs on both servers; `detect_hosts()` determines which direction each sync goes
|
||||
# Drive temperature exit codes respected — skip share or abort all on CRIT.
|
||||
#
|
||||
# 3. Post-sync: DAILY_MAINTENANCE_SCRIPTS (everything except git pull)
|
||||
# media_management.sh → permissions + cleaners + arr cleanup
|
||||
# Media/ + Arrs_Stack/ scripts → permissions, cleaners, arr cleanup, classification
|
||||
# docker_daily_restart.sh → nightly container restarts
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
```
|
||||
@@ -551,12 +557,12 @@ Runs on both servers; `detect_hosts()` determines which direction each sync goes
|
||||
#
|
||||
# HOST1 runs this script at 1am:
|
||||
# → pushes HOST1_DAILY_SYNC_SHARES (Movies, Tv_Shows, Music) → HOST2
|
||||
# → media_management.sh on HOST1's shares
|
||||
# → Media/ + Arrs_Stack/ scripts on HOST1's shares
|
||||
# → docker_daily_restart.sh on HOST1's containers
|
||||
#
|
||||
# HOST2 runs this script at 1am:
|
||||
# → pushes HOST2_DAILY_SYNC_SHARES (Anime_Shows, Anime_Movies) → HOST1
|
||||
# → media_management.sh on HOST2's shares
|
||||
# → Media/ + Arrs_Stack/ scripts on HOST2's shares
|
||||
# → docker_daily_restart.sh on HOST2's containers
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
```
|
||||
@@ -576,9 +582,9 @@ DAILY_MAINTENANCE_SCRIPTS=(
|
||||
"Media/media_shares_permissions.sh" # POST-SYNC — permissions before arr
|
||||
"Media/media_cleaner.sh anime" # POST-SYNC — junk before orphan scan
|
||||
"Media/media_cleaner.sh media" # POST-SYNC
|
||||
"Media/lidarr_cleanup.sh" # POST-SYNC — orphan cleanup last
|
||||
"Media/sonarr_cleanup.sh" # POST-SYNC
|
||||
"Media/radarr_cleanup.sh" # POST-SYNC
|
||||
"Arrs_Stack/lidarr_cleanup.sh" # POST-SYNC — orphan cleanup last
|
||||
"Arrs_Stack/sonarr_cleanup.sh" # POST-SYNC
|
||||
"Arrs_Stack/radarr_cleanup.sh" # POST-SYNC
|
||||
"Docker_Essentials/docker_daily_restart.sh" # POST-SYNC — restarts after everything
|
||||
)
|
||||
|
||||
@@ -618,9 +624,9 @@ DAILY_MAINTENANCE_SCRIPTS=(
|
||||
"Media/media_cleaner.sh anime"
|
||||
"Media/media_cleaner.sh media"
|
||||
"Media/my_new_script.sh" # ← add here in the correct order
|
||||
"Media/lidarr_cleanup.sh"
|
||||
"Media/sonarr_cleanup.sh"
|
||||
"Media/radarr_cleanup.sh"
|
||||
"Arrs_Stack/lidarr_cleanup.sh"
|
||||
"Arrs_Stack/sonarr_cleanup.sh"
|
||||
"Arrs_Stack/radarr_cleanup.sh"
|
||||
"Docker_Essentials/docker_daily_restart.sh"
|
||||
)
|
||||
|
||||
@@ -630,7 +636,7 @@ DAILY_MAINTENANCE_SCRIPTS=(
|
||||
"Media/media_shares_permissions.sh"
|
||||
# "Media/media_cleaner.sh anime" # ← temporarily disabled
|
||||
"Media/media_cleaner.sh media"
|
||||
"Media/lidarr_cleanup.sh"
|
||||
"Arrs_Stack/lidarr_cleanup.sh"
|
||||
...
|
||||
)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -656,7 +662,7 @@ DAILY_MAINTENANCE_SCRIPTS=(
|
||||
|
||||
---
|
||||
|
||||
### ── Relationship to Failover Writeback ──────────────────────────────────────
|
||||
### ── Relationship to Fallback Writeback ──────────────────────────────────────
|
||||
|
||||
```bash
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
@@ -666,7 +672,7 @@ DAILY_MAINTENANCE_SCRIPTS=(
|
||||
# Normal (daily_sync_maintenance.sh):
|
||||
# HOST1 → pushes Movies, Tv_Shows → HOST2
|
||||
#
|
||||
# Tier 4 failover writeback (HOST1 returns after 24hr+ outage):
|
||||
# Tier 4 fallback writeback (HOST1 returns after 24hr+ outage):
|
||||
# HOST2 → pushes Movies, Tv_Shows → HOST1
|
||||
# (HOST2 was running HOST1's arrs and accumulated content)
|
||||
#
|
||||
@@ -708,7 +714,7 @@ maintenance block before the 7am coffee report.
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Containers stop BEFORE sync — clean static source, full bandwidth.
|
||||
# Containers start AFTER sync — on fresh data, in dependency order.
|
||||
# DDNS and failover continue running throughout — only managed containers stop.
|
||||
# DDNS and fallback continue running throughout — only managed containers stop.
|
||||
#
|
||||
# 1. Pre-flight checks — connectivity, remote Docker daemon, remote rootfs
|
||||
# 2. Stop local containers — Emby + auth stack stopped on this server
|
||||
@@ -728,11 +734,11 @@ maintenance block before the 7am coffee report.
|
||||
|
||||
```bash
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Two Emby syncs run in parallel — dirty and clean:
|
||||
# Two Emby sync mechanisms keep the mirror current:
|
||||
#
|
||||
# emby-fallback dirty sync (every 30 minutes via critical_sync_maintenance.sh, Emby running):
|
||||
# watch states, library deltas, user activity — continuous coverage
|
||||
# WAL files excluded — safe to copy while Emby writes
|
||||
# play_state_sync (every 30 minutes via critical_sync_maintenance.sh):
|
||||
# API-based — syncs watched state and resume points directly via Emby API
|
||||
# No rsync of live files — no WAL risk, no partial-write corruption
|
||||
# HOST2 always within 30 minutes of HOST1 on playback state
|
||||
#
|
||||
# weekly clean sync (Sunday 2:30am, Emby stopped):
|
||||
@@ -848,7 +854,7 @@ long tests — that should only run on stable systems that have been up for at l
|
||||
|
||||
Two gates must both pass before any job runs:
|
||||
1. Server uptime ≥ `MONTHLY_UPTIME_THRESHOLD_DAYS`
|
||||
2. Last run ≥ `MONTHLY_RUN_INTERVAL_DAYS` ago (state file on `/boot/config/` — survives reboots)
|
||||
2. Last run ≥ `MONTHLY_RUN_INTERVAL_DAYS` ago (state file in `$STATE_DIR` — survives reboots)
|
||||
|
||||
If either gate fails, the script exits 0 with no output. This is expected — it runs
|
||||
daily and most days are no-ops.
|
||||
@@ -865,7 +871,7 @@ MONTHLY_MAINTENANCE_SCRIPTS=(
|
||||
)
|
||||
MONTHLY_UPTIME_THRESHOLD_DAYS=30
|
||||
MONTHLY_RUN_INTERVAL_DAYS=30
|
||||
MONTHLY_LAST_RUN_FILE="/boot/config/monthly_maintenance_last_run.db"
|
||||
MONTHLY_LAST_RUN_FILE="$STATE_DIR/monthly_maintenance_last_run.db"
|
||||
```
|
||||
|
||||
Scripts are commented out by default — uncomment what applies to your hardware.
|
||||
@@ -882,80 +888,33 @@ monthly_maintenance.sh --log # verbose output from each child script
|
||||
|
||||
---
|
||||
|
||||
## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
## 🧹 media_management.sh
|
||||
## ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
Runs all media maintenance scripts sequentially in the order defined by
|
||||
`MEDIA_MAINTENANCE_JOBS` in `master.conf`. Called by `daily_sync_maintenance.sh`
|
||||
as a post-sync job — not scheduled separately. Available for manual runs when
|
||||
media maintenance is needed outside the normal window.
|
||||
|
||||
```bash
|
||||
# Called by: daily_sync_maintenance.sh (post-sync)
|
||||
# Manual use: run directly for ad hoc media maintenance
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### ── Execution Order ──────────────────────────────────────────────────────────
|
||||
|
||||
```bash
|
||||
# master.conf
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Order is critical — see README-Media.md for detailed explanation.
|
||||
# Short version: permissions first, cleaners second, arr cleanup last.
|
||||
# Each step depends on the previous having completed correctly.
|
||||
#
|
||||
MEDIA_MAINTENANCE_JOBS=(
|
||||
"Media/media_shares_permissions.sh" # 1. permissions — arr cleanup depends on this
|
||||
"Media/media_cleaner.sh anime" # 2. junk removal — orphan scan depends on this
|
||||
"Media/media_cleaner.sh media" # 3. same for media shares
|
||||
"Media/lidarr_cleanup.sh" # 4. orphan cleanup — last, after permissions + clean
|
||||
"Media/sonarr_cleanup.sh" # 5.
|
||||
"Media/radarr_cleanup.sh" # 6.
|
||||
)
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Add a script: insert in the correct position for your use case.
|
||||
# Remove a script: comment it out with #
|
||||
# No changes to media_management.sh needed in either case.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### ── Usage ───────────────────────────────────────────────────────────────────
|
||||
|
||||
```bash
|
||||
media_management.sh # normal run — all jobs in order
|
||||
media_management.sh --dry-run # preview without any deletions or changes
|
||||
media_management.sh --log # verbose output from all jobs
|
||||
media_management.sh --status # show configured job list and exit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ━━━ COMPLETE SCHEDULE ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
```bash
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# At Startup of Array — single entry:
|
||||
# Array start event (Varaverk disks_mounted hook):
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
array_started.sh
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Frequent — every 7 minutes:
|
||||
# Every 7 minutes:
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
*/7 * * * * transcode_management.sh
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Every 30 minutes — auth stack + Emby dirty sync + partnership check:
|
||||
# Every 15 minutes — watchdog cycle:
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
*/15 * * * * watchdog_orchestrator.sh
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Every 30 minutes — auth stack + play_state_sync + partnership check:
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
*/30 * * * * critical_sync_maintenance.sh
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Every 6 hours:
|
||||
# Every 4 hours — arr library sync + failed import recovery + optional rsync:
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
0 */6 * * * arrs_failed_stalled_recovery.sh
|
||||
0 */4 * * * intermediate_sync_maintenance.sh
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Daily — 1am:
|
||||
@@ -970,6 +929,12 @@ array_started.sh
|
||||
50 2 * * 0 CA Auto Update plugin # plugin updates
|
||||
55 2 * * 0 CA container updates # container image updates
|
||||
0 3 * * 0 Network reboot # router/switch restart
|
||||
0 7 * * 0 sunday_morning_coffee_report.sh
|
||||
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# 15th of each month (uptime-gated — silent no-op if uptime < 30 days):
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
0 0 15 * * monthly_maintenance.sh
|
||||
```
|
||||
|
||||
---
|
||||
@@ -977,41 +942,40 @@ array_started.sh
|
||||
## ━━━ ADDING A NEW ORCHESTRATOR ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
If you find yourself running 3+ related scripts on the same schedule, wrap them
|
||||
in a new orchestrator. Model directly on `media_management.sh` which has the
|
||||
in a new orchestrator. Model directly on `daily_sync_maintenance.sh` which has the
|
||||
complete pattern — dry-run passthrough, status display, pass/fail tracking, summary.
|
||||
|
||||
Child-script execution goes through the shared `run_orch_child()` helper in
|
||||
`common.sh` — never hand-roll a per-file `run_job()` loop. It resolves the entry
|
||||
against `$ECOSYSTEM_ROOT`, threads `--dry-run`/`--log` from `$DRY_RUN`/`$ENABLE_LOGGING`
|
||||
automatically (never `$VERBOSE` — nothing in this codebase assigns it), and tracks
|
||||
into `JOB_PASS`/`JOB_FAIL` arrays the caller declares.
|
||||
|
||||
```bash
|
||||
# Minimal skeleton — the full pattern in its simplest form:
|
||||
#!/bin/bash
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
source "$ECOSYSTEM_ROOT/load_config.sh"
|
||||
parse_args "$@"
|
||||
|
||||
SCRIPTS_ROOT="$SCRIPT_DIR/.."
|
||||
PASS=()
|
||||
FAIL=()
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
|
||||
# Read job list from master.conf — never hardcode jobs in the orchestrator
|
||||
for script_entry in "${MY_MAINTENANCE_JOBS[@]:-}"; do
|
||||
[[ -z "$script_entry" ]] && continue
|
||||
|
||||
read -r -a parts <<< "$script_entry"
|
||||
script_path="$SCRIPTS_ROOT/${parts[0]}"
|
||||
script_name=$(basename "${parts[0]}")
|
||||
extra_args=("${parts[@]:1}")
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && extra_args+=("--dry-run")
|
||||
|
||||
if bash "$script_path" "${extra_args[@]}"; then
|
||||
PASS+=("$script_name")
|
||||
else
|
||||
FAIL+=("$script_name")
|
||||
fi
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
|
||||
# One summary — one notification
|
||||
echo "Passed: ${#PASS[@]} Failed: ${#FAIL[@]}"
|
||||
[[ ${#FAIL[@]} -gt 0 ]] && \
|
||||
notify "My maintenance failed on $(hostname) ($MY_ID) — ${FAIL[*]}" \
|
||||
# One summary — one notification. On a frequent (sub-daily) cadence, keep the
|
||||
# healthy path to a single line and reserve the full breakdown for failure/--log —
|
||||
# see watchdog_orchestrator.sh or transcode_management.sh for that split.
|
||||
echo "Passed: ${#JOB_PASS[@]} Failed: ${#JOB_FAIL[@]}"
|
||||
if [[ ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
notify "My maintenance failed on $(hostname) ($MY_ID) — ${JOB_FAIL[*]}" \
|
||||
"My Orchestrator" "warning"
|
||||
exit 1
|
||||
fi
|
||||
exit 0
|
||||
```
|
||||
Regular → Executable
+123
-67
@@ -2,59 +2,111 @@
|
||||
# ==============================================================================================
|
||||
# ================================= Array Start Orchestrator ===================================
|
||||
# ==============================================================================================
|
||||
# Single entry point for "At Startup of Array" in the User Scripts plugin.
|
||||
# Launches everything configured in ARRAY_START_SCRIPTS in master.conf.
|
||||
# This script exits after launching all scripts — unRAID sees it complete normally.
|
||||
#
|
||||
# ── WHAT IT LAUNCHES ──────────────────────────────────────────────────────────────────────────
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Single entry point for array start — fired by the Varaverk plugin's
|
||||
# disks_mounted event hook (Plugin/unraid/event/disks_mounted/array_start_jobs).
|
||||
# Launches everything configured in ARRAY_START_SCRIPTS in master.conf.
|
||||
# This script exits after launching all scripts — the event hook sees it complete normally.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Configured in master.conf ARRAY_START_SCRIPTS — no changes to this script ever needed.
|
||||
# Current order (order matters — see below):
|
||||
#
|
||||
# ONE-SHOT (run and exit naturally):
|
||||
# unRAID_Essentials/inotify_tuning.sh — raise inotify limits before containers start
|
||||
# unRAID_Essentials/docker_syslog_filter.sh — suppress veth log noise before logs fill
|
||||
# unRAID_Essentials/php_fpm_max_children.sh — WebGUI performance tuning
|
||||
# Transcodes/ramdisk_setup.sh — create tmpfs + symlink before Emby starts
|
||||
# System_Essentials/unraid_api_key_renew.sh — re-register Varaverk API key at boot
|
||||
# System_Essentials/inotify_tuning.sh — raise inotify limits before containers start
|
||||
# System_Essentials/docker_syslog_filter.sh — suppress veth log noise before logs fill
|
||||
# System_Essentials/php_fpm_max_children.sh — WebGUI performance tuning
|
||||
# Transcodes/ramdisk_setup.sh — create tmpfs + symlink before Emby starts
|
||||
# Docker_Essentials/docker_network_connect.sh — ensure networks + container connections
|
||||
#
|
||||
# CONTINUOUS (run until array stops):
|
||||
# Watchdogs/stability_watchdog.sh — system health monitor (last line of defense))
|
||||
# Watchdogs/docker_watchdog.sh — container health monitor
|
||||
# Fallback/fallback.sh — mutual failover monitor
|
||||
# Fallback/fallback.sh — mutual fallback monitor
|
||||
#
|
||||
# ── WHY ORDER MATTERS ─────────────────────────────────────────────────────────────────────────
|
||||
# inotify_tuning.sh — must run BEFORE Code-Server and other containers start
|
||||
# containers that start with low inotify limits keep them ✅
|
||||
# docker_syslog_filter — must run BEFORE any container starts creating veth interfaces
|
||||
# ramdisk_setup.sh — must run BEFORE Emby starts transcoding
|
||||
# docker_network_connect — must run BEFORE watchdogs check container states
|
||||
# system_watchdog.sh — before docker_watchdog (system > container priority)
|
||||
# docker_watchdog.sh — before failover (containers must be healthy for failover)
|
||||
# fallback.sh — last — needs everything else stable to make decisions
|
||||
# NOTE: watchdogs (docker, system, stability) are NOT launched here.
|
||||
# They run via watchdog_orchestrator.sh every 15 min (cron), not as daemons.
|
||||
#
|
||||
# ── ONE-SHOT vs CONTINUOUS DETECTION ─────────────────────────────────────────────────────────
|
||||
# Script is launched in background with bash script.sh &
|
||||
# After 1 second: if PID still alive → continuous (running in background)
|
||||
# if PID dead + exit 0 → one-shot completed successfully
|
||||
# if PID dead + exit N → failure
|
||||
# if PID dead + exit 0 → one-shot completed successfully
|
||||
# if PID dead + exit N → failure
|
||||
#
|
||||
# ── SAFEGUARDS ────────────────────────────────────────────────────────────────────────────────
|
||||
# Root check — all launched scripts require root
|
||||
# acquire_lock — prevents duplicate array start launches
|
||||
# detect_hosts() — MY_ID in notifications
|
||||
# validate_unraid_cmd — notify validated before use
|
||||
# chmod +x auto-fix — non-executable scripts fixed before launch
|
||||
# Full path on failure — shows exact path for debugging
|
||||
# notify on failures — alert if any script fails to launch
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Order Is Load-Bearing
|
||||
# inotify limits must be raised before containers start — containers that
|
||||
# start with low limits keep them. Ramdisk must exist before Emby starts.
|
||||
# Docker networks must be connected before watchdogs check container states.
|
||||
# fallback.sh goes last — it needs everything else stable to make decisions.
|
||||
#
|
||||
# Configuration Owns the List
|
||||
# ARRAY_START_SCRIPTS in master.conf is the only place scripts are added or
|
||||
# removed. This orchestrator never needs to be edited to change what runs —
|
||||
# one-shot vs continuous behaviour is auto-detected from the PID after launch.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Every script launched here requires root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents duplicate array start launches. Unraid can fire the array
|
||||
# start hook more than once, and a second pass would re-launch continuous scripts
|
||||
# that are already running.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID for notifications.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if ARRAY_START_SCRIPTS is empty. An empty
|
||||
# list would silently bring the array up with no ramdisk, no network setup, no
|
||||
# watchdogs and no fallback — while reporting a clean start.
|
||||
#
|
||||
# Executable Auto-Fix
|
||||
# Non-executable scripts are chmod +x'd before launch. A permission bit lost to a
|
||||
# git checkout or a file copy should not silently disable a boot-time component.
|
||||
#
|
||||
# Full Path on Failure
|
||||
# Failures report the exact resolved path, so a missing script is immediately
|
||||
# distinguishable from a script that ran and failed.
|
||||
#
|
||||
# Failure Notification
|
||||
# Any script that fails to launch raises a notification — array start is unattended,
|
||||
# so a silent failure here would only surface much later as a missing service.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# ARRAY_START_SCRIPTS — ordered list of scripts to launch at array start
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# array_started.sh — normal launch (called by User Scripts at array start)
|
||||
# array_started.sh --dry-run — show what would be launched without launching
|
||||
# array_started.sh --status — show configured scripts and their current state
|
||||
# array_started.sh --log — verbose output per script
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# array_started.sh
|
||||
# Normal launch — called by Varaverk disks_mounted event hook.
|
||||
#
|
||||
# array_started.sh --dry-run
|
||||
# Show what would be launched without launching.
|
||||
#
|
||||
# array_started.sh --status
|
||||
# Show configured scripts and their current state.
|
||||
#
|
||||
# array_started.sh --log
|
||||
# Verbose output per script.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
@@ -72,15 +124,21 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#ARRAY_START_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "ARRAY_START_SCRIPTS is empty — no array start scripts will run"
|
||||
error "Check ARRAY_START_SCRIPTS in master.conf"
|
||||
notify "array start scripts skipped on $(hostname) ($MY_ID) — ARRAY_START_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — scripts will not be launched"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -121,6 +179,13 @@ fi
|
||||
# ==============================================================================================
|
||||
# ━━━ Launch Scripts ━━━
|
||||
# ==============================================================================================
|
||||
if [[ ${#ARRAY_START_SCRIPTS[@]} -eq 0 ]]; then
|
||||
notify "Array started on $LOCAL_SERVER_NAME ($MY_ID) but ARRAY_START_SCRIPTS is empty — boot sequence skipped. Check master.conf." \
|
||||
"Array Start" "alert"
|
||||
error "ARRAY_START_SCRIPTS is empty — check master.conf"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR Array Start — $MY_ID — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
log "Ecosystem root: $ECOSYSTEM_ROOT"
|
||||
@@ -128,9 +193,8 @@ log "Launching ${#ARRAY_START_SCRIPTS[@]} script(s)..."
|
||||
echo ""
|
||||
|
||||
START=$(date +%s)
|
||||
LAUNCHED=0
|
||||
FAILED=0
|
||||
FAILED_SCRIPTS=()
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
|
||||
for relative_path in "${ARRAY_START_SCRIPTS[@]}"; do
|
||||
[[ -z "$relative_path" ]] && continue
|
||||
@@ -142,8 +206,7 @@ for relative_path in "${ARRAY_START_SCRIPTS[@]}"; do
|
||||
if [[ ! -f "$SCRIPT_PATH" ]]; then
|
||||
error "$SCRIPT_NAME — not found"
|
||||
error " Expected: $SCRIPT_PATH"
|
||||
(( FAILED++ ))
|
||||
FAILED_SCRIPTS+=("$SCRIPT_NAME")
|
||||
JOB_FAIL+=("$SCRIPT_NAME")
|
||||
continue
|
||||
fi
|
||||
|
||||
@@ -152,15 +215,14 @@ for relative_path in "${ARRAY_START_SCRIPTS[@]}"; do
|
||||
warn "$SCRIPT_NAME — not executable, fixing..."
|
||||
chmod +x "$SCRIPT_PATH" || {
|
||||
error "$SCRIPT_NAME — chmod +x failed"
|
||||
(( FAILED++ ))
|
||||
FAILED_SCRIPTS+=("$SCRIPT_NAME")
|
||||
JOB_FAIL+=("$SCRIPT_NAME")
|
||||
continue
|
||||
}
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would launch: $SCRIPT_NAME"
|
||||
(( LAUNCHED++ ))
|
||||
JOB_PASS+=("$SCRIPT_NAME")
|
||||
continue
|
||||
fi
|
||||
|
||||
@@ -173,20 +235,19 @@ for relative_path in "${ARRAY_START_SCRIPTS[@]}"; do
|
||||
|
||||
if kill -0 "$PID" 2>/dev/null; then
|
||||
# Still running → continuous script
|
||||
warn "$SCRIPT_NAME — running (PID $PID) ✅"
|
||||
(( LAUNCHED++ ))
|
||||
log "$SCRIPT_NAME — running (PID $PID) ✅"
|
||||
JOB_PASS+=("$SCRIPT_NAME")
|
||||
else
|
||||
# Exited — check if one-shot success or failure
|
||||
wait "$PID"
|
||||
EXIT_CODE=$?
|
||||
if [[ "$EXIT_CODE" -eq 0 ]]; then
|
||||
log "$SCRIPT_NAME — completed (one-shot) ✅"
|
||||
(( LAUNCHED++ ))
|
||||
echo "$SCRIPT_NAME — completed (one-shot) ✅"
|
||||
JOB_PASS+=("$SCRIPT_NAME")
|
||||
else
|
||||
error "$SCRIPT_NAME — exited with code $EXIT_CODE"
|
||||
error " Path: $SCRIPT_PATH"
|
||||
(( FAILED++ ))
|
||||
FAILED_SCRIPTS+=("$SCRIPT_NAME")
|
||||
JOB_FAIL+=("$SCRIPT_NAME")
|
||||
fi
|
||||
fi
|
||||
|
||||
@@ -200,18 +261,13 @@ END=$(date +%s)
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY ARRAY START SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_SUCCESS Launched: $LAUNCHED"
|
||||
[[ "$FAILED" -gt 0 ]] && echo "$ICON_ERROR Failed: $FAILED — ${FAILED_SCRIPTS[*]}"
|
||||
echo "$ICON_SUCCESS Launched: ${#JOB_PASS[@]}"
|
||||
[[ ${#JOB_FAIL[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${#JOB_FAIL[@]} — ${JOB_FAIL[*]}"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no scripts launched"
|
||||
elif [[ "$FAILED" -gt 0 ]]; then
|
||||
warn "Status: $FAILED script(s) failed — ${FAILED_SCRIPTS[*]}"
|
||||
notify "Array start on $(hostname) ($MY_ID) — $FAILED script(s) failed: ${FAILED_SCRIPTS[*]}" \
|
||||
"Array Start" "warning"
|
||||
else
|
||||
echo "$ICON_DONE Status: all $LAUNCHED script(s) launched ✅"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
# The configured list is the denominator — a script the conf names but that never launched is
|
||||
# skipped, not absent, and only shows up if something counts it.
|
||||
JOB_COUNT="${#ARRAY_START_SCRIPTS[@]}"
|
||||
orchestrator_summary "ARRAY START" "$START" "Array Start"
|
||||
exit $?
|
||||
@@ -2,40 +2,91 @@
|
||||
# ==============================================================================================
|
||||
# ================================= Array Stop Orchestrator ====================================
|
||||
# ==============================================================================================
|
||||
# Planned shutdown orchestrator — stops all active processes cleanly before array maintenance.
|
||||
# Runs ARRAY_STOP_SCRIPTS from master.conf sequentially, each confirmed complete before next.
|
||||
#
|
||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Planned shutdown orchestrator — stops all active processes cleanly before
|
||||
# array maintenance. Runs ARRAY_STOP_SCRIPTS from master.conf sequentially,
|
||||
# each confirmed complete before the next starts.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# 1. user_scripts_stop.sh — kill background user scripts (prevents new operations)
|
||||
# 2. rsync_stop.sh --rsync-only — kill rsync; skip container recovery (handled in step 4)
|
||||
# 3. mover_stop.sh — stop mover after rsync (both write to same paths)
|
||||
# 4. docker_container_stop.sh — stop all containers one-by-one with verification
|
||||
#
|
||||
# ── WHY THIS ORDER ────────────────────────────────────────────────────────────────────────────
|
||||
# User scripts stopped first — they can spawn new rsync/docker operations mid-shutdown.
|
||||
# Rsync before mover — both write to the same paths; running together risks corruption.
|
||||
# Containers last — apps should stay available as long as possible during shutdown prep.
|
||||
# Unlike array_started.sh, all scripts run in the foreground. Each must complete
|
||||
# (pass or fail) before the next starts — a failed stop is noted but does not
|
||||
# prevent remaining steps from running.
|
||||
#
|
||||
# ── SEQUENTIAL vs BACKGROUND ─────────────────────────────────────────────────────────────────
|
||||
# Unlike array_started.sh, all scripts run in the foreground. Each must complete (pass or fail)
|
||||
# before the next starts — a failed stop is noted but does not prevent remaining steps.
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── SAFEGUARDS ────────────────────────────────────────────────────────────────────────────────
|
||||
# Root check — all stop scripts require root
|
||||
# acquire_lock — prevents concurrent array stop runs
|
||||
# detect_hosts() — MY_ID in notifications and logs
|
||||
# validate_unraid_cmd — notify validated before use
|
||||
# Non-fatal steps — a failed step is logged but remaining steps still run
|
||||
# notify on failures — alert if any stop script fails
|
||||
# Order Is Load-Bearing
|
||||
# User scripts are stopped first — they can spawn new rsync or docker operations
|
||||
# mid-shutdown. Rsync stops before mover — both write to the same paths and
|
||||
# running together risks corruption. Containers stop last — apps should stay
|
||||
# available as long as possible during shutdown prep.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# A failed stop step is logged and notified but does not abort the sequence.
|
||||
# Remaining scripts still run — a partial stop is better than a halted one.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Every stop script launched here requires root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents concurrent array stop runs. Two overlapping shutdown
|
||||
# sequences would fight over the same containers.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID for notifications and logs.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if ARRAY_STOP_SCRIPTS is empty. An empty
|
||||
# list means the array stops without saving the conf cache or gracefully stopping
|
||||
# containers — the failure would only be discovered at the next boot.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# A failing stop script is recorded and the remaining ones still run. Abandoning the
|
||||
# shutdown sequence partway would leave more state unsaved than continuing does.
|
||||
#
|
||||
# Failure Notification
|
||||
# Any failing stop script raises a notification. Shutdown is unattended and its
|
||||
# failures are invisible until they cause a problem on the way back up.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# ARRAY_STOP_SCRIPTS — ordered list of stop scripts to run
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# array_stopping.sh — run full stop sequence
|
||||
# array_stopping.sh --dry-run — preview without stopping anything
|
||||
# array_stopping.sh --status — show configured scripts and exit
|
||||
# array_stopping.sh --log — verbose output
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# array_stopping.sh
|
||||
# Run full stop sequence.
|
||||
#
|
||||
# array_stopping.sh --dry-run
|
||||
# Preview without stopping anything.
|
||||
#
|
||||
# array_stopping.sh --status
|
||||
# Show configured scripts and exit.
|
||||
#
|
||||
# array_stopping.sh --log
|
||||
# Verbose output.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
@@ -53,10 +104,6 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
@@ -67,6 +114,16 @@ fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#ARRAY_STOP_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "ARRAY_STOP_SCRIPTS is empty — no array stop scripts will run"
|
||||
error "Check ARRAY_STOP_SCRIPTS in master.conf"
|
||||
notify "array stop scripts skipped on $(hostname) ($MY_ID) — ARRAY_STOP_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no stop scripts will be executed"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -104,8 +161,8 @@ echo "$ICON_GEAR Running ${#ARRAY_STOP_SCRIPTS[@]} stop script(s) sequentially..
|
||||
echo ""
|
||||
|
||||
START=$(date +%s)
|
||||
PASSED=()
|
||||
FAILED=()
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
STEP=0
|
||||
|
||||
for entry in "${ARRAY_STOP_SCRIPTS[@]}"; do
|
||||
@@ -113,44 +170,11 @@ for entry in "${ARRAY_STOP_SCRIPTS[@]}"; do
|
||||
(( STEP++ ))
|
||||
|
||||
read -r -a parts <<< "$entry"
|
||||
script_path="$ECOSYSTEM_ROOT/${parts[0]}"
|
||||
script_name=$(basename "${parts[0]}")
|
||||
extra_args=("${parts[@]:1}")
|
||||
|
||||
echo "━━━ $ICON_GEAR Step $STEP: $script_name${extra_args:+ ${extra_args[*]}} ━━━"
|
||||
|
||||
if [[ ! -f "$script_path" ]]; then
|
||||
error "$script_name — not found at $script_path"
|
||||
FAILED+=("$script_name")
|
||||
echo ""
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ ! -x "$script_path" ]]; then
|
||||
warn "$script_name — not executable, fixing..."
|
||||
chmod +x "$script_path" || {
|
||||
error "$script_name — chmod +x failed"
|
||||
FAILED+=("$script_name")
|
||||
echo ""
|
||||
continue
|
||||
}
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would run: $script_name ${extra_args[*]}"
|
||||
PASSED+=("$script_name")
|
||||
echo ""
|
||||
continue
|
||||
fi
|
||||
|
||||
if bash "$script_path" "${extra_args[@]}"; then
|
||||
log "$script_name — done ✅"
|
||||
PASSED+=("$script_name")
|
||||
else
|
||||
warn "$script_name — failed (exit $?) — continuing to next step"
|
||||
FAILED+=("$script_name")
|
||||
fi
|
||||
|
||||
run_orch_child "$entry"
|
||||
echo ""
|
||||
done
|
||||
|
||||
@@ -162,22 +186,10 @@ END=$(date +%s)
|
||||
echo "━━━━━ $ICON_SUMMARY ARRAY STOP SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
[[ ${#PASSED[@]} -gt 0 ]] && echo "$ICON_DONE Passed: ${PASSED[*]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
[[ ${#JOB_PASS[@]} -gt 0 ]] && echo "$ICON_DONE Passed: ${JOB_PASS[*]}"
|
||||
[[ ${#JOB_FAIL[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${JOB_FAIL[*]}"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ ${#FAILED[@]} -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: all $STEP step(s) complete ✅"
|
||||
notify "Array stop complete on $(hostname) ($MY_ID) — $STEP step(s) done" \
|
||||
"Array Stop" "normal"
|
||||
else
|
||||
warn "Status: ${#FAILED[@]} step(s) failed — ${FAILED[*]}"
|
||||
notify "Array stop on $(hostname) ($MY_ID) — ${#FAILED[@]} step(s) failed: ${FAILED[*]}" \
|
||||
"Array Stop" "warning"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
JOB_COUNT="$STEP"
|
||||
orchestrator_summary "ARRAY STOP" "$START" "Array Stop"
|
||||
exit $?
|
||||
|
||||
Regular → Executable
+130
-76
@@ -2,50 +2,105 @@
|
||||
# ==============================================================================================
|
||||
# ============================= Critical Sync Maintenance ======================================
|
||||
# ==============================================================================================
|
||||
# Orchestrator for time-sensitive syncs that run every 30 minutes.
|
||||
# Keeps the mirror current between the less frequent daily and weekly windows.
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Orchestrator for time-sensitive syncs running every 30 minutes. Keeps the
|
||||
# mirror current between the less frequent daily and weekly windows.
|
||||
# Schedule: */30 * * * * (every 30 minutes via User Scripts plugin)
|
||||
#
|
||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||
# 1. Critical-Data rsync — auth stack, NPM config, certs (containers stopped both sides)
|
||||
# 2. emby-fallback rsync — dirty Emby sync (watch states, library — Emby stays running)
|
||||
# 3. CRITICAL_MAINTENANCE_SCRIPTS — any scripts configured for critical window
|
||||
# 4. partnership --check — read both state files, detect changes, act accordingly
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── WHY EVERY 30 MINUTES ──────────────────────────────────────────────────────────────────────
|
||||
# Auth stack changes (new users, proxy rules, certs) propagate within 30min ✅
|
||||
# Emby watch states stay in sync — mirror users see correct playback position ✅
|
||||
# Partnership state changes detected and acted on quickly ✅
|
||||
# Lock prevents: daily rsync doing Critical-Data mid-critical window ✅
|
||||
# 1. Critical-Data rsync — auth stack, NPM config, certs (containers stopped both sides)
|
||||
# 2. CRITICAL_MAINTENANCE_SCRIPTS — play_state_sync + any other per-window scripts
|
||||
# 3. partnership --check — read both state files, detect changes, act accordingly
|
||||
#
|
||||
# ── RSYNC GATE ────────────────────────────────────────────────────────────────────────────────
|
||||
# RSYNC GATE
|
||||
# RSYNC_ENABLED=false → skips all syncs (global gate)
|
||||
# CRITICAL_RSYNC_ENABLED=false → skips critical syncs only (per-orchestrator gate)
|
||||
# partnership --check always runs regardless — state check doesn't need rsync
|
||||
# partnership --check always runs regardless — state check doesn't need rsync.
|
||||
#
|
||||
# ── LOCK BEHAVIOUR ────────────────────────────────────────────────────────────────────────────
|
||||
# acquire_lock "strict" — if previous 30min run still going, skip this cycle entirely
|
||||
# Critical-Data taking > 30min is a problem worth knowing about
|
||||
# Strict mode prevents pile-up without waiting — log and move on ✅
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── SILENT WHEN HEALTHY ───────────────────────────────────────────────────────────────────────
|
||||
# Runs 48 times per day — clean runs must produce zero output ✅
|
||||
# Only failures and notable events produce visible output
|
||||
# Silent When Healthy
|
||||
# Runs 48 times per day — clean runs must produce zero output. Only failures
|
||||
# and notable events produce visible output.
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# CRITICAL_RSYNC_ENABLED — enable/disable rsync section
|
||||
# CRITICAL_SYNC_SHARES — shares synced every 30min (HOST*_CRITICAL_SYNC_SHARES)
|
||||
# CRITICAL_MAINTENANCE_SCRIPTS — scripts run in critical window (optional)
|
||||
# PARTNERSHIP_ENABLED — enable/disable partnership check
|
||||
# Auth-First Window
|
||||
# Auth stack changes (new users, proxy rules, certs) propagate within 30min.
|
||||
# Emby watch states stay in sync — mirror users see correct playback position.
|
||||
# Partnership state changes detected and acted on quickly.
|
||||
#
|
||||
# Strict Lock, Never Queue
|
||||
# acquire_lock "strict" — if the previous 30-min run is still going, skip
|
||||
# this cycle entirely. Critical-Data taking > 30min is a problem worth
|
||||
# knowing about. Strict mode prevents pile-up without waiting.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# rsync over SSH and container stop/start both require root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock in strict mode — a cycle is skipped rather than queued if the previous
|
||||
# one is still running. At a 30-minute cadence, queuing would let a slow sync stack
|
||||
# windows behind it.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID and REMOTE_ID for routing and logs.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if CRITICAL_MAINTENANCE_SCRIPTS is empty —
|
||||
# a silently empty critical tier would stop downloader resets and play-state sync
|
||||
# while still reporting success every 30 minutes.
|
||||
#
|
||||
# Remote IP Resolution
|
||||
# resolve_remote_ip confirms the partner is reachable before any transfer is attempted.
|
||||
#
|
||||
# RSYNC_ENABLED Gate
|
||||
# The global kill switch is respected before any rsync call, so disabling rsync
|
||||
# ecosystem-wide genuinely stops it here too.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# A failing job is recorded and the rest of the tier still runs.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# CRITICAL_RSYNC_ENABLED — enable/disable rsync section
|
||||
# CRITICAL_SYNC_SHARES — shares synced every 30min (HOST*_CRITICAL_SYNC_SHARES)
|
||||
# CRITICAL_MAINTENANCE_SCRIPTS — scripts run in critical window (optional)
|
||||
# PARTNERSHIP_ENABLED — enable/disable partnership check
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# critical_sync_maintenance.sh
|
||||
# Normal run.
|
||||
#
|
||||
# critical_sync_maintenance.sh --dry-run
|
||||
# Preview syncs without transferring.
|
||||
#
|
||||
# critical_sync_maintenance.sh --log
|
||||
# Verbose per-share output.
|
||||
#
|
||||
# critical_sync_maintenance.sh --status
|
||||
# Show configuration and exit.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# critical_sync_maintenance.sh — normal run
|
||||
# critical_sync_maintenance.sh --dry-run — preview syncs without transferring
|
||||
# critical_sync_maintenance.sh --log — verbose per-share output
|
||||
# critical_sync_maintenance.sh --status — show configuration and exit
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
@@ -59,14 +114,20 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock "strict"
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#CRITICAL_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "CRITICAL_MAINTENANCE_SCRIPTS is empty — no critical maintenance scripts will run"
|
||||
error "Check CRITICAL_MAINTENANCE_SCRIPTS in master.conf"
|
||||
notify "critical maintenance scripts skipped on $(hostname) ($MY_ID) — CRITICAL_MAINTENANCE_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
resolve_remote_ip
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
||||
@@ -119,6 +180,10 @@ RSYNC_OK=false
|
||||
PASS=()
|
||||
FAIL=()
|
||||
|
||||
log "$ICON_SYNC Critical shares (${#CRITICAL_SYNC_SHARES[@]}): $(for s in "${CRITICAL_SYNC_SHARES[@]}"; do printf '%s ' "$(basename "${s%%|*}")"; done)"
|
||||
[[ ${#CRITICAL_MAINTENANCE_SCRIPTS[@]} -gt 0 ]] && \
|
||||
log "$ICON_GEAR Maintenance scripts: $(for s in "${CRITICAL_MAINTENANCE_SCRIPTS[@]}"; do printf '%s ' "$(basename "${s%% *}")"; done)"
|
||||
|
||||
if ! check_rsync_enabled "CRITICAL"; then
|
||||
echo "Critical rsync disabled — skipping sync, running partnership check only"
|
||||
elif [[ ${#CRITICAL_SYNC_SHARES[@]} -eq 0 ]]; then
|
||||
@@ -153,7 +218,7 @@ else
|
||||
|
||||
if [[ "$RSYNC_EXIT" -eq 0 ]]; then
|
||||
PASS+=("$SHARE_NAME")
|
||||
log "$SHARE_NAME — done in $SHARE_DUR ✅"
|
||||
echo "$SHARE_NAME — done in $SHARE_DUR ✅"
|
||||
RSYNC_OK=true
|
||||
else
|
||||
FAIL+=("$SHARE_NAME")
|
||||
@@ -165,29 +230,12 @@ fi
|
||||
# ==============================================================================================
|
||||
# ━━━ Critical Maintenance Scripts ━━━
|
||||
# ==============================================================================================
|
||||
if [[ ${#CRITICAL_MAINTENANCE_SCRIPTS[@]} -gt 0 ]]; then
|
||||
for script_entry in "${CRITICAL_MAINTENANCE_SCRIPTS[@]}"; do
|
||||
[[ -z "$script_entry" || "$script_entry" == \#* ]] && continue
|
||||
|
||||
SCRIPT_PATH="$SCRIPT_DIR/../${script_entry%% *}"
|
||||
SCRIPT_ARGS="${script_entry#* }"
|
||||
[[ "$SCRIPT_ARGS" == "$script_entry" ]] && SCRIPT_ARGS=""
|
||||
[[ "$DRY_RUN" == true ]] && SCRIPT_ARGS="$SCRIPT_ARGS --dry-run"
|
||||
|
||||
SCRIPT_NAME=$(basename "$SCRIPT_PATH")
|
||||
|
||||
if [[ ! -f "$SCRIPT_PATH" ]]; then
|
||||
warn "$SCRIPT_NAME not found at $SCRIPT_PATH — skipping"
|
||||
continue
|
||||
fi
|
||||
|
||||
log "Running: $SCRIPT_NAME"
|
||||
bash "$SCRIPT_PATH" $SCRIPT_ARGS
|
||||
EXIT_CODE=$?
|
||||
[[ "$EXIT_CODE" -ne 0 ]] && \
|
||||
warn "$SCRIPT_NAME exited with code $EXIT_CODE"
|
||||
done
|
||||
fi
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
for script_entry in "${CRITICAL_MAINTENANCE_SCRIPTS[@]}"; do
|
||||
[[ -z "$script_entry" || "$script_entry" == \#* ]] && continue
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Partnership Check ━━━
|
||||
@@ -196,9 +244,28 @@ if [[ "${PARTNERSHIP_ENABLED:-false}" == true ]]; then
|
||||
PARTNER_DRY=""
|
||||
[[ "$DRY_RUN" == true ]] && PARTNER_DRY="--dry-run"
|
||||
|
||||
# A successful rsync proves the partner answered; a failed one is evidence it did not. But
|
||||
# rsync being switched off is neither — and it used to be read as "unseen", so the offline
|
||||
# counter climbed every 30 minutes toward the 30-day auto-offboard on a partnership whose
|
||||
# only fault was that RSYNC_ENABLED=false. That is how a deliberately paused sync ends up
|
||||
# dismantling the partnership it was paused for. With no rsync attempt there is nothing to
|
||||
# report, so the check runs without touching the counter either way.
|
||||
# Tier 2 counts as "switched off" here exactly as much as Tier 1 does. The guard below used
|
||||
# to test RSYNC_ENABLED alone, but it is CRITICAL_RSYNC_ENABLED that governs whether this
|
||||
# orchestrator attempts an rsync at all — so with Tier 1 open and Tier 2 closed, no transfer
|
||||
# was attempted, RSYNC_OK stayed false, and the run fell through to --remote-unseen and
|
||||
# incremented the counter every 30 minutes against a partner that was answering fine.
|
||||
#
|
||||
# Onboard Step 1d now leaves precisely that posture on purpose — Tier 1 open so provisioning
|
||||
# can run, every Tier 2 gate closed so nothing is scheduled. A freshly onboarded, perfectly
|
||||
# healthy partnership would have auto-offboarded itself 30 days later.
|
||||
if [[ "$RSYNC_OK" == true ]]; then
|
||||
bash "$SCRIPT_DIR/../Partnership/partnership_manager.sh" \
|
||||
--check --remote-seen $PARTNER_DRY
|
||||
elif [[ "${RSYNC_ENABLED:-false}" != true || "${CRITICAL_RSYNC_ENABLED:-false}" != true ]]; then
|
||||
echo "Critical rsync gated off — partnership check runs, offline counter untouched"
|
||||
bash "$SCRIPT_DIR/../Partnership/partnership_manager.sh" \
|
||||
--check $PARTNER_DRY
|
||||
else
|
||||
bash "$SCRIPT_DIR/../Partnership/partnership_manager.sh" \
|
||||
--check --remote-unseen $PARTNER_DRY
|
||||
@@ -213,20 +280,7 @@ fi
|
||||
END=$(date +%s)
|
||||
DURATION=$(format_duration $(( END - START )))
|
||||
|
||||
# Silent when healthy — only show summary if there were failures or notable events
|
||||
if [[ ${#FAIL[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY CRITICAL SYNC SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $DURATION"
|
||||
[[ ${#PASS[@]} -gt 0 ]] && echo "Synced: ${PASS[*]}"
|
||||
echo "$ICON_ERROR Failed: ${FAIL[*]}"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
notify "Critical sync failed on $(hostname) ($MY_ID) — ${FAIL[*]}" \
|
||||
"Critical Sync" "warning"
|
||||
exit 1
|
||||
else
|
||||
echo "Critical sync complete — $MY_ID — ${DURATION} — ${#PASS[@]} share(s)"
|
||||
fi
|
||||
|
||||
exit 0
|
||||
# Standard ending, quiet mode — 30-min cadence, so a healthy cycle stays one line.
|
||||
[[ ${#PASS[@]} -gt 0 ]] && echo "Synced: ${PASS[*]}"
|
||||
orchestrator_summary "CRITICAL SYNC" "$START" "Critical Sync" quiet
|
||||
exit $?
|
||||
Regular → Executable
+132
-109
@@ -2,76 +2,136 @@
|
||||
# ==============================================================================================
|
||||
# ============================= Daily Sync Maintenance =========================================
|
||||
# ==============================================================================================
|
||||
# Daily orchestrator — runs the full daily maintenance window in the correct order.
|
||||
# Schedule: 0 1 * * * (1am daily via User Scripts plugin)
|
||||
#
|
||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||
# Pre-sync:
|
||||
# git_pull_execute.sh — pull latest scripts first, always
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Daily maintenance window orchestrator — runs the full daily sequence in the
|
||||
# correct order. Schedule: 0 1 * * * (1am daily via User Scripts plugin)
|
||||
#
|
||||
# Arr Sync (arr_sync.sh):
|
||||
# Syncs Lidarr/Sonarr/Radarr libraries across all nodes bidirectionally.
|
||||
# All nodes agree on tracked library before any files are transferred.
|
||||
# Remote nodes that don't have an arr running are skipped gracefully.
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Rsync window (DAILY_SYNC_SHARES per host):
|
||||
# HOST*_DAILY_SYNC_SHARES — media shares spread to all nodes (no --delete)
|
||||
# HOST*_PERSONAL_SHARES — encrypted personal shares
|
||||
# Pre-sync:
|
||||
# git_pull_execute.sh — pull latest scripts first, always
|
||||
#
|
||||
# Post-sync maintenance (DAILY_MAINTENANCE_SCRIPTS):
|
||||
# media_shares_permissions.sh — fix ownership before arr cleanup
|
||||
# media_cleaner.sh anime — remove junk from anime shares
|
||||
# media_cleaner.sh media — remove junk from media shares
|
||||
# lidarr_cleanup.sh — remove orphaned music files (local arr = truth)
|
||||
# sonarr_cleanup.sh — remove orphaned TV files (local arr = truth)
|
||||
# radarr_cleanup.sh — remove orphaned movie files (local arr = truth)
|
||||
# docker_daily_restart.sh — restart containers needing daily restart
|
||||
# Arr Sync (arr_sync.sh):
|
||||
# Syncs Lidarr/Sonarr/Radarr libraries across all nodes bidirectionally.
|
||||
# All nodes agree on tracked library before any files are transferred.
|
||||
# Remote nodes that don't have an arr running are skipped gracefully.
|
||||
#
|
||||
# ── WHY ORDER MATTERS ─────────────────────────────────────────────────────────────────────────
|
||||
# git pull first — maintenance runs on latest code, not yesterday's
|
||||
# arr sync before rsync — all nodes track the same library before files are spread;
|
||||
# prevents remote arrs from searching for content already owned
|
||||
# rsync before cleanup — cleanup sees fully spread state, rsync has no --delete
|
||||
# permissions before arr cleanup — arrs need correct ownership to delete/rename
|
||||
# arr cleanup after permissions — clean ownership = successful orphan deletion
|
||||
# docker restart last — containers already processed by cleanup
|
||||
# Rsync window (DAILY_SYNC_SHARES per host):
|
||||
# HOST*_DAILY_SYNC_SHARES — media shares spread to all nodes (no --delete)
|
||||
# HOST*_PERSONAL_SHARES — encrypted personal shares
|
||||
#
|
||||
# ── HOST AWARENESS ────────────────────────────────────────────────────────────────────────────
|
||||
# Bidirectional — same script runs on both servers, correct direction automatic.
|
||||
# detect_hosts() aliases DAILY_SYNC_SHARES and PERSONAL_SHARES from HOST*_ vars.
|
||||
# No manual HOST1/HOST2 comparisons — MY_ID routes correctly on any server.
|
||||
# Post-sync maintenance (DAILY_MAINTENANCE_SCRIPTS):
|
||||
# media_shares_permissions.sh — fix ownership before arr cleanup
|
||||
# media_cleaner.sh anime/media — remove junk from media and anime shares
|
||||
# lidarr/sonarr/radarr_cleanup.sh — remove orphaned files (local arr = truth)
|
||||
# docker_daily_restart.sh — restart containers needing daily restart
|
||||
#
|
||||
# ── DRIVE TEMP HANDLING ───────────────────────────────────────────────────────────────────────
|
||||
# rsync.sh returns exit codes for temperature issues:
|
||||
# exit 1 = temp WARN — skip this share, continue to next
|
||||
# exit 2 = temp CRITICAL — abort ALL remaining syncs in this window
|
||||
# All other failures — skip share, continue to next
|
||||
# DRIVE TEMP HANDLING (rsync.sh exit codes):
|
||||
# exit 1 = temp WARN → skip this share, continue to next
|
||||
# exit 2 = temp CRITICAL → abort ALL remaining syncs in this window
|
||||
#
|
||||
# ── SILENT WHEN HEALTHY ───────────────────────────────────────────────────────────────────────
|
||||
# Runs daily at 1am — clean run should produce minimal output.
|
||||
# Each job reports log() on success (silent), warn()/error() on failure (visible).
|
||||
# Summary always shown — gives window timing and share/job counts.
|
||||
# Notify only on failure — successful daily maintenance doesn't need notification.
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── CONFIGURATION (master.conf + host*.conf) ───────────────────────────────────────────
|
||||
# HOST*_DAILY_SYNC_SHARES — shares pushed to mirror each day
|
||||
# HOST*_PERSONAL_SHARES — encrypted personal shares
|
||||
# DAILY_MAINTENANCE_SCRIPTS — maintenance jobs (permissions, cleanup, restart)
|
||||
# DAILY_RSYNC_ENABLED — enable/disable rsync section
|
||||
# Order Is Load-Bearing
|
||||
# git pull first — maintenance runs on latest code, not yesterday's. Arr sync
|
||||
# before rsync — all nodes track the same library before files are spread,
|
||||
# preventing remote arrs from searching for content already owned. Rsync before
|
||||
# cleanup — cleanup sees fully spread state. Permissions before arr cleanup —
|
||||
# arrs need correct ownership to delete/rename. Docker restart last — containers
|
||||
# already processed by cleanup.
|
||||
#
|
||||
# Host-Aware Routing
|
||||
# Bidirectional — same script runs on both servers, correct direction automatic.
|
||||
# detect_hosts() aliases DAILY_SYNC_SHARES and PERSONAL_SHARES from HOST*_ vars.
|
||||
# No manual HOST1/HOST2 comparisons needed.
|
||||
#
|
||||
# Silent When Healthy
|
||||
# Runs daily at 1am — clean runs produce minimal output. Each job logs silently
|
||||
# on success; failures surface to warn()/error(). Notify only on failure.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# rsync and docker operations require root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents two daily windows overlapping — the window is long and a
|
||||
# second pass would contend for the same shares and containers.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases the correct per-host share and script lists.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if DAILY_MAINTENANCE_SCRIPTS is empty. This
|
||||
# is the largest tier in the ecosystem — an empty list would silently skip git pull,
|
||||
# permissions, cleaners, arr cleanup and docker updates while reporting a clean run.
|
||||
#
|
||||
# Connectivity Check
|
||||
# check_connectivity is verified before any rsync is attempted.
|
||||
#
|
||||
# Remote Rootfs Check
|
||||
# check_remote_rootfs aborts rsync if the remote rootfs is nearly full, rather than
|
||||
# pushing data to a partner that cannot hold it.
|
||||
#
|
||||
# Drive Temperature Escalation
|
||||
# rsync.sh's exit code is honoured per share: exit 1 (temp WARN) skips that share and
|
||||
# continues; exit 2 (temp CRITICAL) sets ABORT_ALL_SYNCS so every remaining share in
|
||||
# the window is skipped and a notification is raised. Continuing to hammer drives that
|
||||
# are already too hot is how a thermal warning becomes a dead disk.
|
||||
#
|
||||
# Non-Fatal Jobs
|
||||
# A failed job is logged and the remaining jobs still run. Partial completion of a
|
||||
# maintenance window beats abandoning it at the first error.
|
||||
#
|
||||
# Quiet on Success
|
||||
# A successful daily run produces no notification — only failures surface.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_DAILY_SYNC_SHARES — shares pushed to mirror each day
|
||||
# HOST*_PERSONAL_SHARES — encrypted personal shares
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# DAILY_MAINTENANCE_SCRIPTS — maintenance jobs (permissions, cleanup, restart)
|
||||
# DAILY_RSYNC_ENABLED — enable/disable rsync section
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# daily_sync_maintenance.sh
|
||||
# Normal run.
|
||||
#
|
||||
# daily_sync_maintenance.sh --dry-run
|
||||
# Preview without syncing or changing.
|
||||
#
|
||||
# daily_sync_maintenance.sh --log
|
||||
# Verbose per-share/per-job output.
|
||||
#
|
||||
# daily_sync_maintenance.sh --status
|
||||
# Show configured shares and jobs.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# daily_sync_maintenance.sh — normal run
|
||||
# daily_sync_maintenance.sh --dry-run — preview without syncing or changing
|
||||
# daily_sync_maintenance.sh --log — verbose per-share/per-job output
|
||||
# daily_sync_maintenance.sh --status — show configured shares and jobs
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
RSYNC_SCRIPT="$SCRIPT_DIR/../Rsync/rsync.sh"
|
||||
SCRIPTS_ROOT="$SCRIPT_DIR/.."
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
RSYNC_SCRIPT="$ECOSYSTEM_ROOT/Rsync/rsync.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
@@ -83,10 +143,6 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
if ! command -v docker &>/dev/null; then
|
||||
error "Docker command not found"
|
||||
@@ -94,6 +150,16 @@ if ! command -v docker &>/dev/null; then
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#DAILY_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "DAILY_MAINTENANCE_SCRIPTS is empty — no daily maintenance scripts will run"
|
||||
error "Check DAILY_MAINTENANCE_SCRIPTS in master.conf"
|
||||
notify "daily maintenance scripts skipped on $(hostname) ($MY_ID) — DAILY_MAINTENANCE_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
resolve_remote_ip
|
||||
|
||||
acquire_lock
|
||||
@@ -155,35 +221,6 @@ for script_entry in "${DAILY_MAINTENANCE_SCRIPTS[@]}"; do
|
||||
fi
|
||||
done
|
||||
|
||||
# Helper — run a maintenance script, track pass/fail
|
||||
run_job() {
|
||||
local script_entry="$1"
|
||||
local extra_dry=""
|
||||
[[ "$DRY_RUN" == true ]] && extra_dry="--dry-run"
|
||||
|
||||
read -r -a script_args <<< "$script_entry"
|
||||
local script_path="$SCRIPTS_ROOT/${script_args[0]}"
|
||||
local script_name
|
||||
script_name=$(basename "${script_args[0]}")
|
||||
local extra_args=("${script_args[@]:1}")
|
||||
|
||||
if [[ ! -f "$script_path" ]]; then
|
||||
error "$script_name — not found at $script_path"
|
||||
JOB_FAIL+=("$script_name")
|
||||
return 1
|
||||
fi
|
||||
|
||||
log "Running: $script_name ${extra_args[*]}"
|
||||
# shellcheck disable=SC2086
|
||||
if bash "$script_path" "${extra_args[@]}" $extra_dry; then
|
||||
log "$script_name — done ✅"
|
||||
JOB_PASS+=("$script_name ${extra_args[*]}")
|
||||
else
|
||||
error "$script_name — failed (exit $?)"
|
||||
JOB_FAIL+=("$script_name ${extra_args[*]}")
|
||||
fi
|
||||
}
|
||||
|
||||
WINDOW_START=$(date +%s)
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
@@ -201,7 +238,7 @@ if [[ ${#PRE_SYNC_SCRIPTS[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
echo "━━━ $ICON_GIT Pre-sync ━━━"
|
||||
for script_entry in "${PRE_SYNC_SCRIPTS[@]}"; do
|
||||
run_job "$script_entry"
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
fi
|
||||
|
||||
@@ -211,7 +248,7 @@ fi
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Arr Sync ━━━"
|
||||
|
||||
ARR_SYNC_SCRIPT="$SCRIPTS_ROOT/Media/arr_sync.sh"
|
||||
ARR_SYNC_SCRIPT="$ECOSYSTEM_ROOT/Arrs_Stack/arr_sync.sh"
|
||||
if [[ "${ARR_SYNC_ENABLED:-true}" != "true" ]]; then
|
||||
echo "ARR_SYNC_ENABLED=false — skipping"
|
||||
elif [[ ! -f "$ARR_SYNC_SCRIPT" ]]; then
|
||||
@@ -275,7 +312,7 @@ else
|
||||
case "$RSYNC_EXIT" in
|
||||
0)
|
||||
PASS+=("$SHARE_NAME")
|
||||
log "$SHARE_NAME — done ✅"
|
||||
echo "$SHARE_NAME — done ✅"
|
||||
;;
|
||||
1)
|
||||
FAIL+=("$SHARE_NAME:temp-warn")
|
||||
@@ -307,7 +344,7 @@ if [[ ${#POST_SYNC_SCRIPTS[@]} -gt 0 ]]; then
|
||||
echo "━━━ $ICON_CLEAN Post-sync Maintenance ━━━"
|
||||
for script_entry in "${POST_SYNC_SCRIPTS[@]}"; do
|
||||
echo ""
|
||||
run_job "$script_entry"
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
fi
|
||||
|
||||
@@ -317,11 +354,6 @@ WINDOW_END=$(date +%s)
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY DAILY MAINTENANCE SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Window: $(date -d @"$WINDOW_START" '+%Y-%m-%d %H:%M:%S') → $(date -d @"$WINDOW_END" '+%H:%M:%S')"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( WINDOW_END - WINDOW_START )))"
|
||||
echo ""
|
||||
|
||||
echo "$ICON_SYNC Shares ($SHARE_COUNT):"
|
||||
for entry in "${SHARE_TIMES[@]}"; do
|
||||
@@ -345,16 +377,7 @@ if [[ ${#JOB_PASS[@]} -gt 0 || ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
fi
|
||||
|
||||
TOTAL_FAIL=$(( ${#FAIL[@]} + ${#JOB_FAIL[@]} ))
|
||||
|
||||
if [[ "$TOTAL_FAIL" -gt 0 ]]; then
|
||||
warn "Status: $TOTAL_FAIL failure(s)"
|
||||
notify "Daily maintenance completed with failures on $(hostname) ($MY_ID) — shares: ${#FAIL[@]}/$SHARE_COUNT failed, jobs: ${#JOB_FAIL[@]} failed" \
|
||||
"Daily Maintenance" "warning"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 1
|
||||
else
|
||||
echo "$ICON_DONE Status: all complete — ${#PASS[@]} share(s) synced, ${#JOB_PASS[@]} job(s) run"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
# Standard ending — derives skipped from SHARE_COUNT, so a run with rsync gated off reports
|
||||
# PARTIAL instead of "all complete".
|
||||
orchestrator_summary "DAILY MAINTENANCE" "$WINDOW_START" "Daily Maintenance"
|
||||
exit $?
|
||||
Regular → Executable
+171
-119
@@ -2,63 +2,114 @@
|
||||
# ==============================================================================================
|
||||
# =========================== Intermediate Sync Maintenance ====================================
|
||||
# ==============================================================================================
|
||||
# 4-hour orchestrator — arr library reconciliation, artwork fetching, and optional rsync.
|
||||
# Schedule: 0 */4 * * * (every 4 hours)
|
||||
#
|
||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||
# 1. arr_sync.sh — sync Lidarr/Sonarr/Radarr libraries across all nodes
|
||||
# 2. Rsync window (optional) — INTERMEDIATE_SYNC_SHARES, if any configured
|
||||
# 3. INTERMEDIATE_MAINTENANCE_SCRIPTS — artwork fetch and any future 4-hour jobs
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# 4-hour orchestrator — arr library reconciliation, artwork fetching, and
|
||||
# optional rsync. Schedule: 0 */4 * * * (every 4 hours)
|
||||
#
|
||||
# ── WHY A SEPARATE ORCHESTRATOR ───────────────────────────────────────────────────────────────
|
||||
# arr libraries need to converge more frequently than once a day. If a remote node adds
|
||||
# something at 2am, the next daily window is 23 hours away — remote arrs search for content
|
||||
# they don't know is already owned. Running every 4 hours closes that gap.
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# lidarr_missing_art.sh is idempotent — skips existing files, runs fast after initial fill.
|
||||
# Pairing it here means artwork catches up within 4 hours of a new album landing.
|
||||
# 1. conf_sync.sh — refresh partner conf cache in RAM (/tmp/varaverk/conf/),
|
||||
# both directions: pull theirs, push ours
|
||||
# 2. arr_sync.sh — sync Lidarr/Sonarr/Radarr libraries across all nodes
|
||||
# 3. Rsync window (optional) — INTERMEDIATE_SYNC_SHARES, if any configured
|
||||
# 4. INTERMEDIATE_MAINTENANCE_SCRIPTS — artwork fetch and any future 4-hour jobs
|
||||
#
|
||||
# Rsync is optional — INTERMEDIATE_SYNC_SHARES empty by default. Add shares to the config
|
||||
# if a subset of data needs mid-day propagation (e.g. watch state, metadata). Full media
|
||||
# share sync stays in the daily window.
|
||||
# DRIVE TEMP HANDLING (rsync.sh exit codes):
|
||||
# exit 1 = temp WARN → skip this share, continue to next
|
||||
# exit 2 = temp CRITICAL → abort ALL remaining syncs in this window
|
||||
#
|
||||
# ── HOST AWARENESS ────────────────────────────────────────────────────────────────────────────
|
||||
# detect_hosts() sets MY_ID and aliases HOST*_INTERMEDIATE_SYNC_SHARES → INTERMEDIATE_SYNC_SHARES.
|
||||
# Each server can have a different set of mid-day shares — configure in host*.conf.
|
||||
# Each script in INTERMEDIATE_MAINTENANCE_SCRIPTS handles its own host logic.
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── DRIVE TEMP HANDLING ───────────────────────────────────────────────────────────────────────
|
||||
# Same as daily_sync_maintenance.sh:
|
||||
# exit 1 = temp WARN — skip this share, continue to next
|
||||
# exit 2 = temp CRITICAL — abort ALL remaining syncs in this window
|
||||
# Closes the Library Gap
|
||||
# Arr libraries need to converge more frequently than once a day. If a remote
|
||||
# node adds something at 2am, the next daily window is 23 hours away — remote
|
||||
# arrs search for content they don't know is already owned. Running every 4
|
||||
# hours closes that gap.
|
||||
#
|
||||
# ── SAFEGUARDS ────────────────────────────────────────────────────────────────────────────────
|
||||
# Root check — scripts called here require root
|
||||
# acquire_lock — prevents concurrent intermediate windows
|
||||
# check_connectivity — verified before any rsync (skipped if no shares)
|
||||
# check_remote_rootfs — aborts rsync if remote rootfs nearly full
|
||||
# Non-fatal jobs — a failed arr_sync warns but does not block rsync or artwork fetch
|
||||
# Silent on success — runs 4x/day, only failures warrant notification
|
||||
# Optional Rsync Layer
|
||||
# INTERMEDIATE_SYNC_SHARES is empty by default — the rsync step is skipped
|
||||
# entirely when nothing is configured. Add shares only if a subset of data
|
||||
# needs mid-day propagation. Full media share sync stays in the daily window.
|
||||
#
|
||||
# ── CONFIGURATION ─────────────────────────────────────────────────────────────────────────────
|
||||
# host*.conf: HOST*_INTERMEDIATE_SYNC_SHARES — shares synced mid-day (empty = rsync skipped)
|
||||
# master.conf: INTERMEDIATE_RSYNC_ENABLED — enable/disable rsync section (default: true)
|
||||
# master.conf: INTERMEDIATE_MAINTENANCE_SCRIPTS — jobs run after rsync
|
||||
# master.conf: ARR_SYNC_ENABLED — toggle inside arr_sync.sh
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Every script called from here requires root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents concurrent intermediate windows.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases the correct per-host share lists.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if INTERMEDIATE_MAINTENANCE_SCRIPTS is empty,
|
||||
# rather than running six no-op windows a day that all report success.
|
||||
#
|
||||
# Connectivity Check
|
||||
# check_connectivity is verified before any rsync, and skipped entirely when no shares
|
||||
# are configured — there is nothing to reach a partner for.
|
||||
#
|
||||
# Remote Rootfs Check
|
||||
# check_remote_rootfs aborts rsync if the remote rootfs is nearly full.
|
||||
#
|
||||
# Drive Temperature Escalation
|
||||
# rsync.sh's exit code is honoured per share: exit 1 skips that share, exit 2 aborts
|
||||
# every remaining sync in the window and notifies.
|
||||
#
|
||||
# Non-Fatal Jobs
|
||||
# A failing arr_sync warns but does not block rsync or the artwork fetch that follow it.
|
||||
#
|
||||
# Minimal on Success
|
||||
# Runs six times a day, so the full breakdown only prints on failure or with --log.
|
||||
# A quiet run is the normal outcome and should not fill the log.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# host*.conf
|
||||
#
|
||||
# HOST*_INTERMEDIATE_SYNC_SHARES — shares synced mid-day (empty = rsync skipped)
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# INTERMEDIATE_RSYNC_ENABLED — enable/disable rsync section (default: true)
|
||||
# INTERMEDIATE_MAINTENANCE_SCRIPTS — jobs run after rsync
|
||||
# ARR_SYNC_ENABLED — toggle inside arr_sync.sh
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# intermediate_sync_maintenance.sh
|
||||
# Normal run.
|
||||
#
|
||||
# intermediate_sync_maintenance.sh --dry-run
|
||||
# Preview without changes.
|
||||
#
|
||||
# intermediate_sync_maintenance.sh --log
|
||||
# Verbose per-job output.
|
||||
#
|
||||
# intermediate_sync_maintenance.sh --status
|
||||
# Show configured shares/jobs and exit.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# intermediate_sync_maintenance.sh — normal run
|
||||
# intermediate_sync_maintenance.sh --dry-run — preview without changes
|
||||
# intermediate_sync_maintenance.sh --log — verbose per-job output
|
||||
# intermediate_sync_maintenance.sh --status — show configured shares/jobs and exit
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
RSYNC_SCRIPT="$SCRIPT_DIR/../Rsync/rsync.sh"
|
||||
SCRIPTS_ROOT="$SCRIPT_DIR/.."
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
RSYNC_SCRIPT="$ECOSYSTEM_ROOT/Rsync/rsync.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
@@ -70,47 +121,24 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#INTERMEDIATE_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "INTERMEDIATE_MAINTENANCE_SCRIPTS is empty — no intermediate maintenance scripts will run"
|
||||
error "Check INTERMEDIATE_MAINTENANCE_SCRIPTS in master.conf"
|
||||
notify "intermediate maintenance scripts skipped on $(hostname) ($MY_ID) — INTERMEDIATE_MAINTENANCE_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
resolve_remote_ip
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no changes will be made"
|
||||
|
||||
# ── Helper — run a maintenance job, track pass/fail ───────────────────────────────────────────
|
||||
run_job() {
|
||||
local script_entry="$1"
|
||||
local extra_dry=""
|
||||
[[ "$DRY_RUN" == true ]] && extra_dry="--dry-run"
|
||||
|
||||
read -r -a script_args <<< "$script_entry"
|
||||
local script_path="$SCRIPTS_ROOT/${script_args[0]}"
|
||||
local script_name
|
||||
script_name=$(basename "${script_args[0]}")
|
||||
local extra_args=("${script_args[@]:1}")
|
||||
|
||||
if [[ ! -f "$script_path" ]]; then
|
||||
error "$script_name — not found at $script_path"
|
||||
JOB_FAIL+=("$script_name")
|
||||
return 1
|
||||
fi
|
||||
|
||||
log "Running: $script_name ${extra_args[*]}"
|
||||
# shellcheck disable=SC2086
|
||||
if bash "$script_path" "${extra_args[@]}" $extra_dry; then
|
||||
log "$script_name — done ✅"
|
||||
JOB_PASS+=("$script_name ${extra_args[*]}")
|
||||
else
|
||||
warn "$script_name — failed (exit $?)"
|
||||
JOB_FAIL+=("$script_name ${extra_args[*]}")
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
@@ -154,13 +182,39 @@ SHARE_TIMES=()
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR Intermediate Sync — $MY_ID — $(date '+%Y-%m-%d %H:%M:%S') ━━━"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Conf Sync ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR Conf Sync ━━━"
|
||||
|
||||
# Full sync, not --pull-only. The push half was written as an event-driven fast path for the
|
||||
# conf-save hook, but no such hook was ever built — so outside array start nothing pushed this
|
||||
# host's conf to its partners at all, and a partner's copy of our conf stayed at whatever it was
|
||||
# when we last rebooted. Pull alone kept our view of them fresh while their view of us decayed.
|
||||
CONF_SYNC_SCRIPT="$ECOSYSTEM_ROOT/System_Essentials/conf_sync.sh"
|
||||
if [[ ! -f "$CONF_SYNC_SCRIPT" ]]; then
|
||||
warn "conf_sync.sh not found — skipping partner conf refresh"
|
||||
else
|
||||
_conf_args=()
|
||||
[[ "$DRY_RUN" == true ]] && _conf_args+=("--dry-run")
|
||||
if bash "$CONF_SYNC_SCRIPT" "${_conf_args[@]}"; then
|
||||
echo "Partner conf cache refreshed ✅"
|
||||
JOB_PASS+=("conf_sync.sh")
|
||||
else
|
||||
warn "Partner conf sync failed — cache may be stale"
|
||||
JOB_FAIL+=("conf_sync.sh")
|
||||
fi
|
||||
unset _conf_args
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Arr Sync ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━ $ICON_SYNC Arr Sync ━━━"
|
||||
|
||||
ARR_SYNC_SCRIPT="$SCRIPTS_ROOT/Media/arr_sync.sh"
|
||||
ARR_SYNC_SCRIPT="$ECOSYSTEM_ROOT/Arrs_Stack/arr_sync.sh"
|
||||
if [[ "${ARR_SYNC_ENABLED:-true}" != "true" ]]; then
|
||||
echo "ARR_SYNC_ENABLED=false — skipping"
|
||||
elif [[ ! -f "$ARR_SYNC_SCRIPT" ]]; then
|
||||
@@ -225,7 +279,7 @@ else
|
||||
case "$RSYNC_EXIT" in
|
||||
0)
|
||||
PASS+=("$SHARE_NAME")
|
||||
log "$SHARE_NAME — done ✅"
|
||||
echo "$SHARE_NAME — done ✅"
|
||||
;;
|
||||
1)
|
||||
FAIL+=("$SHARE_NAME:temp-warn")
|
||||
@@ -257,7 +311,7 @@ if [[ ${#INTERMEDIATE_MAINTENANCE_SCRIPTS[@]} -gt 0 ]]; then
|
||||
for script_entry in "${INTERMEDIATE_MAINTENANCE_SCRIPTS[@]}"; do
|
||||
[[ -z "$script_entry" ]] && continue
|
||||
echo ""
|
||||
run_job "$script_entry"
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
fi
|
||||
|
||||
@@ -266,47 +320,45 @@ WINDOW_END=$(date +%s)
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY INTERMEDIATE SYNC SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Window: $(date -d @"$WINDOW_START" '+%Y-%m-%d %H:%M:%S') → $(date -d @"$WINDOW_END" '+%H:%M:%S')"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( WINDOW_END - WINDOW_START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$SHARE_COUNT" -gt 0 ]]; then
|
||||
echo "$ICON_SYNC Shares ($SHARE_COUNT):"
|
||||
for entry in "${SHARE_TIMES[@]}"; do
|
||||
sname="${entry%%:*}"
|
||||
sdur="${entry##*:}"
|
||||
if printf '%s\n' "${FAIL[@]}" | grep -q "^${sname}"; then
|
||||
echo " $ICON_ERROR $sname — $(format_duration "$sdur")"
|
||||
else
|
||||
echo " $ICON_DONE $sname — $(format_duration "$sdur")"
|
||||
fi
|
||||
done
|
||||
echo " Passed: ${#PASS[@]}/$SHARE_COUNT Failed: ${#FAIL[@]}/$SHARE_COUNT"
|
||||
echo ""
|
||||
fi
|
||||
|
||||
if [[ ${#JOB_PASS[@]} -gt 0 || ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
echo "$ICON_GEAR Jobs:"
|
||||
for job in "${JOB_PASS[@]}"; do echo " $ICON_DONE $job"; done
|
||||
for job in "${JOB_FAIL[@]}"; do echo " $ICON_ERROR $job"; done
|
||||
echo ""
|
||||
fi
|
||||
|
||||
TOTAL_FAIL=$(( ${#FAIL[@]} + ${#JOB_FAIL[@]} ))
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ "$TOTAL_FAIL" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: all complete ✅ — ${#JOB_PASS[@]} job(s) run, ${#PASS[@]}/$SHARE_COUNT share(s) synced"
|
||||
else
|
||||
warn "Status: $TOTAL_FAIL failure(s)"
|
||||
notify "Intermediate sync failed on $(hostname) ($MY_ID) — shares: ${#FAIL[@]}/$SHARE_COUNT failed, jobs: ${#JOB_FAIL[@]} failed" \
|
||||
"Intermediate Sync" "warning"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
# Full breakdown on failure or --log; minimal one-liner otherwise (4-hour cadence — keep it quiet).
|
||||
SHOW_FULL=false
|
||||
[[ "$TOTAL_FAIL" -gt 0 || "$ENABLE_LOGGING" == true ]] && SHOW_FULL=true
|
||||
|
||||
[[ "$TOTAL_FAIL" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
if [[ "$SHOW_FULL" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY INTERMEDIATE SYNC SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Window: $(date -d @"$WINDOW_START" '+%Y-%m-%d %H:%M:%S') → $(date -d @"$WINDOW_END" '+%H:%M:%S')"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( WINDOW_END - WINDOW_START )))"
|
||||
echo ""
|
||||
|
||||
if [[ "$SHARE_COUNT" -gt 0 ]]; then
|
||||
echo "$ICON_SYNC Shares ($SHARE_COUNT):"
|
||||
for entry in "${SHARE_TIMES[@]}"; do
|
||||
sname="${entry%%:*}"
|
||||
sdur="${entry##*:}"
|
||||
if printf '%s\n' "${FAIL[@]}" | grep -q "^${sname}"; then
|
||||
echo " $ICON_ERROR $sname — $(format_duration "$sdur")"
|
||||
else
|
||||
echo " $ICON_DONE $sname — $(format_duration "$sdur")"
|
||||
fi
|
||||
done
|
||||
echo " Passed: ${#PASS[@]}/$SHARE_COUNT Failed: ${#FAIL[@]}/$SHARE_COUNT"
|
||||
echo ""
|
||||
fi
|
||||
|
||||
if [[ ${#JOB_PASS[@]} -gt 0 || ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
echo "$ICON_GEAR Jobs:"
|
||||
for job in "${JOB_PASS[@]}"; do echo " $ICON_DONE $job"; done
|
||||
for job in "${JOB_FAIL[@]}"; do echo " $ICON_ERROR $job"; done
|
||||
echo ""
|
||||
fi
|
||||
fi
|
||||
|
||||
# Standard ending, quiet mode — 4-hour cadence, so an OK cycle is one parseable line and
|
||||
# anything skipped or failed expands to the full block on its own.
|
||||
_mode=quiet; [[ "$ENABLE_LOGGING" == true ]] && _mode=full
|
||||
orchestrator_summary "INTERMEDIATE SYNC" "$WINDOW_START" "Intermediate Sync" "$_mode"
|
||||
exit $?
|
||||
|
||||
Regular → Executable
+114
-81
@@ -2,39 +2,108 @@
|
||||
# ==============================================================================================
|
||||
# ========================= Monthly Maintenance Orchestrator ===================================
|
||||
# ==============================================================================================
|
||||
# Uptime-triggered monthly maintenance — runs heavy tasks that need a stable, settled system.
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Uptime-triggered monthly maintenance — runs heavy tasks that need a stable,
|
||||
# settled system. Schedule: 0 0 15 * * (15th of each month at midnight)
|
||||
# Fires only when BOTH gates pass:
|
||||
# 1. Server uptime >= MONTHLY_UPTIME_THRESHOLD_DAYS days
|
||||
# 2. Last run was >= MONTHLY_RUN_INTERVAL_DAYS days ago (or never run)
|
||||
#
|
||||
# ── WHY UPTIME-GATED ─────────────────────────────────────────────────────────────────────────
|
||||
# A scheduled reboot resets uptime. Monthly tasks (ZFS scrub, SMART long test) need a
|
||||
# stable, settled system — not one that just rebooted. Uptime-gating ensures maintenance
|
||||
# only runs after the server has been healthy for a full month, never immediately post-boot.
|
||||
# If uptime or interval gate is not met on the 15th, the run is skipped until next month.
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── HOW TO CALL ──────────────────────────────────────────────────────────────────────────────
|
||||
# Schedule: 0 0 15 * * (15th of each month at midnight)
|
||||
# Silent exit 0 when either gate is not met. Only outputs when maintenance actually fires.
|
||||
# Runs MONTHLY_MAINTENANCE_SCRIPTS sequentially when both gates pass.
|
||||
# Silent exit 0 when either gate is not met — only outputs when maintenance fires.
|
||||
# If uptime or interval gate is not met on the 15th, the run is skipped until
|
||||
# next month.
|
||||
#
|
||||
# ── STATE FILE ────────────────────────────────────────────────────────────────────────────────
|
||||
# MONTHLY_LAST_RUN_FILE — /boot/config — survives reboots, available before array starts.
|
||||
# Written after each run (pass or partial fail). Format: Unix timestamp.
|
||||
# A reboot does NOT reset the last-run state — the interval gate survives independently
|
||||
# of the uptime gate. Both must pass before maintenance fires again.
|
||||
# STATE FILE
|
||||
# MONTHLY_LAST_RUN_FILE lives on /boot/config — survives reboots, available
|
||||
# before the array starts. Written after each run (pass or partial fail).
|
||||
# Format: Unix timestamp. A reboot does NOT reset the last-run state — the
|
||||
# interval gate survives independently of the uptime gate.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Uptime Gate Ensures Stability
|
||||
# A scheduled reboot resets uptime. Monthly tasks (ZFS scrub, SMART long test)
|
||||
# need a stable, settled system — not one that just rebooted. Both gates must
|
||||
# pass before maintenance fires, ensuring the server has been healthy for a
|
||||
# full month.
|
||||
#
|
||||
# State Survives Reboots
|
||||
# MONTHLY_LAST_RUN_FILE is on /boot/config (USB flash), not on the array.
|
||||
# It is always available regardless of array state, so the interval gate is
|
||||
# never lost to a reboot.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# ZFS scrub and SMART tests require root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents concurrent monthly runs. These are long jobs — a scrub can run
|
||||
# for hours — and two at once would double the I/O cost for no benefit.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID for notifications and logs.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if MONTHLY_MAINTENANCE_SCRIPTS is empty. A
|
||||
# monthly job that silently does nothing is the hardest kind to notice missing.
|
||||
#
|
||||
# Uptime Gate
|
||||
# MONTHLY_UPTIME_THRESHOLD_DAYS must be met before the run proceeds. Heavy full-disk
|
||||
# work immediately after a boot competes with everything else still starting up.
|
||||
#
|
||||
# Interval Gate
|
||||
# MONTHLY_RUN_INTERVAL_DAYS since the last successful run must have elapsed. The
|
||||
# schedule fires more often than the work should actually happen, so the gate — not
|
||||
# the cron entry — is what defines the real cadence.
|
||||
#
|
||||
# Force Override
|
||||
# --force bypasses both gates for a deliberate manual run.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# A failing job is recorded and the remaining jobs still run.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# MONTHLY_MAINTENANCE_SCRIPTS — ordered list of scripts to run
|
||||
# MONTHLY_UPTIME_THRESHOLD_DAYS — minimum uptime in days before maintenance fires
|
||||
# MONTHLY_RUN_INTERVAL_DAYS — minimum days since last run before running again
|
||||
# MONTHLY_LAST_RUN_FILE — state file path — /boot/config, survives reboots
|
||||
# MONTHLY_LAST_RUN_FILE — state file path (/boot/config — survives reboots)
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# monthly_maintenance.sh
|
||||
# Normal run — uptime + interval gates enforced.
|
||||
#
|
||||
# monthly_maintenance.sh --dry-run
|
||||
# Preview gate state and scripts without running.
|
||||
#
|
||||
# monthly_maintenance.sh --status
|
||||
# Show gate state, last run, and configured scripts.
|
||||
#
|
||||
# monthly_maintenance.sh --force
|
||||
# Bypass uptime + interval gates (manual override).
|
||||
#
|
||||
# monthly_maintenance.sh --log
|
||||
# Verbose output.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# monthly_maintenance.sh — normal run (uptime + interval gates enforced)
|
||||
# monthly_maintenance.sh --dry-run — preview gate state and scripts without running
|
||||
# monthly_maintenance.sh --status — show gate state, last run, and configured scripts
|
||||
# monthly_maintenance.sh --force — bypass uptime + interval gates (manual override)
|
||||
# monthly_maintenance.sh --log — verbose output
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
@@ -66,13 +135,23 @@ acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#MONTHLY_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "MONTHLY_MAINTENANCE_SCRIPTS is empty — no monthly maintenance scripts will run"
|
||||
error "Check MONTHLY_MAINTENANCE_SCRIPTS in master.conf"
|
||||
notify "monthly maintenance scripts skipped on $(hostname) ($MY_ID) — MONTHLY_MAINTENANCE_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no scripts will be executed"
|
||||
[[ "$FORCE_RUN" == true ]] && warn "FORCE — uptime and interval gates bypassed"
|
||||
|
||||
# Defaults — overridden by master.conf values
|
||||
MONTHLY_UPTIME_THRESHOLD_DAYS="${MONTHLY_UPTIME_THRESHOLD_DAYS:-30}"
|
||||
MONTHLY_RUN_INTERVAL_DAYS="${MONTHLY_RUN_INTERVAL_DAYS:-30}"
|
||||
MONTHLY_LAST_RUN_FILE="${MONTHLY_LAST_RUN_FILE:-/boot/config/monthly_maintenance_last_run.db}"
|
||||
MONTHLY_LAST_RUN_FILE="${MONTHLY_LAST_RUN_FILE:-${STATE_DIR:-/tmp}/monthly_maintenance_last_run.db}"
|
||||
|
||||
UPTIME_THRESHOLD_SECS=$(( MONTHLY_UPTIME_THRESHOLD_DAYS * 86400 ))
|
||||
INTERVAL_SECS=$(( MONTHLY_RUN_INTERVAL_DAYS * 86400 ))
|
||||
@@ -163,11 +242,11 @@ fi
|
||||
# ==============================================================================================
|
||||
if [[ "$FORCE_RUN" != true ]]; then
|
||||
if [[ "$UPTIME_GATE_PASS" != true ]]; then
|
||||
log "Uptime gate — ${_uptime_days}d ${_uptime_hrs}h / ${MONTHLY_UPTIME_THRESHOLD_DAYS}d — not yet due"
|
||||
echo "Uptime gate not met — ${_uptime_days}d ${_uptime_hrs}h / ${MONTHLY_UPTIME_THRESHOLD_DAYS}d — no-op"
|
||||
exit 0
|
||||
fi
|
||||
if [[ "$INTERVAL_GATE_PASS" != true ]]; then
|
||||
log "Interval gate — last run ${DAYS_SINCE_LAST_RUN} ago / ${MONTHLY_RUN_INTERVAL_DAYS}d — not yet due"
|
||||
echo "Interval gate not met — last run ${DAYS_SINCE_LAST_RUN} ago / ${MONTHLY_RUN_INTERVAL_DAYS}d — no-op"
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
@@ -187,8 +266,8 @@ echo "$ICON_GEAR Scripts: ${#MONTHLY_MAINTENANCE_SCRIPTS[@]}"
|
||||
echo ""
|
||||
|
||||
START=$(date +%s)
|
||||
PASSED=()
|
||||
FAILED=()
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
STEP=0
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -199,47 +278,11 @@ for entry in "${MONTHLY_MAINTENANCE_SCRIPTS[@]}"; do
|
||||
(( STEP++ ))
|
||||
|
||||
read -r -a parts <<< "$entry"
|
||||
script_path="$ECOSYSTEM_ROOT/${parts[0]}"
|
||||
script_name=$(basename "${parts[0]}")
|
||||
extra_args=("${parts[@]:1}")
|
||||
|
||||
echo "━━━ $ICON_GEAR Step $STEP: $script_name${extra_args:+ ${extra_args[*]}} ━━━"
|
||||
|
||||
if [[ ! -f "$script_path" ]]; then
|
||||
error "$script_name — not found at $script_path"
|
||||
FAILED+=("$script_name")
|
||||
echo ""
|
||||
continue
|
||||
fi
|
||||
|
||||
if [[ ! -x "$script_path" ]]; then
|
||||
warn "$script_name — not executable, fixing..."
|
||||
chmod +x "$script_path" || {
|
||||
error "$script_name — chmod +x failed"
|
||||
FAILED+=("$script_name")
|
||||
echo ""
|
||||
continue
|
||||
}
|
||||
fi
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — would run: $script_name${extra_args:+ ${extra_args[*]}}"
|
||||
PASSED+=("$script_name")
|
||||
echo ""
|
||||
continue
|
||||
fi
|
||||
|
||||
local_args=()
|
||||
[[ "$VERBOSE" == true ]] && local_args+=("--log")
|
||||
|
||||
if bash "$script_path" "${extra_args[@]}" "${local_args[@]}"; then
|
||||
log "$script_name — done ✅"
|
||||
PASSED+=("$script_name")
|
||||
else
|
||||
warn "$script_name — failed (exit $?) — continuing to next step"
|
||||
FAILED+=("$script_name")
|
||||
fi
|
||||
|
||||
run_orch_child "$entry"
|
||||
echo ""
|
||||
done
|
||||
|
||||
@@ -260,22 +303,12 @@ fi
|
||||
echo "━━━━━ $ICON_SUMMARY MONTHLY MAINTENANCE SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( END - START )))"
|
||||
[[ ${#PASSED[@]} -gt 0 ]] && echo "$ICON_DONE Passed: ${PASSED[*]}"
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${FAILED[*]}"
|
||||
[[ ${#JOB_PASS[@]} -gt 0 ]] && echo "$ICON_DONE Passed: ${JOB_PASS[*]}"
|
||||
[[ ${#JOB_FAIL[@]} -gt 0 ]] && echo "$ICON_ERROR Failed: ${JOB_FAIL[*]}"
|
||||
echo ""
|
||||
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ ${#FAILED[@]} -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: all $STEP step(s) complete ✅"
|
||||
notify "Monthly maintenance complete on $(hostname) ($MY_ID) — $STEP step(s) done" \
|
||||
"Monthly Maintenance" "normal"
|
||||
else
|
||||
warn "Status: ${#FAILED[@]} step(s) failed — ${FAILED[*]}"
|
||||
notify "Monthly maintenance on $(hostname) ($MY_ID) — ${#FAILED[@]} step(s) failed: ${FAILED[*]}" \
|
||||
"Monthly Maintenance" "warning"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
[[ ${#FAILED[@]} -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
# STEP is what this orchestrator expected to run, so it is the denominator that makes a skipped
|
||||
# step visible rather than absent.
|
||||
JOB_COUNT="$STEP"
|
||||
orchestrator_summary "MONTHLY MAINTENANCE" "$START" "Monthly Maintenance"
|
||||
exit $?
|
||||
|
||||
@@ -2,70 +2,135 @@
|
||||
# ==============================================================================================
|
||||
# ============================= Sunday Morning Coffee Report ===================================
|
||||
# ==============================================================================================
|
||||
#
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Weekly monitoring orchestrator — runs all Sunday monitor scripts in sequence.
|
||||
# Designed to be read over coffee Sunday morning while the system is fully caught up
|
||||
# from the 2:30am maintenance window.
|
||||
# Schedule: 0 7 * * 0 (7am Sunday — after weekly_sync_maintenance.sh finishes at ~3am)
|
||||
# Designed to be read over coffee while the system is fully caught up from the
|
||||
# 2:30am maintenance window.
|
||||
# Schedule: 0 7 * * 0 (7am Sunday — after weekly_sync_maintenance.sh finishes)
|
||||
#
|
||||
# ── SCRIPTS (master.conf COFFEE_REPORT_SCRIPTS) ───────────────────────────────────────────────
|
||||
# Each script runs independently, logs to its own output, and notifies on findings.
|
||||
# All scripts receive --dry-run and --log flags from this orchestrator when set.
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── HOST AWARENESS ────────────────────────────────────────────────────────────────────────────
|
||||
# detect_hosts() sets MY_ID — used in banner and summary.
|
||||
# HOST1 primary: runs all scripts. HOST2: limited to host-aware scripts only.
|
||||
# Runs COFFEE_REPORT_SCRIPTS from master.conf in order. Each script runs
|
||||
# independently, produces its own output, and notifies on findings.
|
||||
# All scripts receive --dry-run and --log flags from this orchestrator when set.
|
||||
#
|
||||
# HOST1 primary: runs all scripts.
|
||||
# HOST2: limited to host-aware scripts only — each script handles its own host logic.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Independent, Isolated Scripts
|
||||
# Each monitor script is fully self-contained — a failure in one does not
|
||||
# prevent the others from running. The orchestrator logs the failure and
|
||||
# continues to the next script.
|
||||
#
|
||||
# Runs After the Full Weekly Window
|
||||
# Scheduled 4+ hours after weekly_sync_maintenance.sh — the system is fully
|
||||
# synced and containers are back up before any monitoring reads run. Reports
|
||||
# reflect the settled post-maintenance state.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Required for the docker and system reads the report is built from.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents overlapping weekly runs.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID for the banner and summary.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if COFFEE_REPORT_SCRIPTS is empty. A report
|
||||
# that silently contains nothing still arrives looking like a report.
|
||||
#
|
||||
# Read-Only by Composition
|
||||
# Every child here is a reporting script. This orchestrator changes nothing itself —
|
||||
# it only sequences reads and assembles their output.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# A failing script is logged and the remaining ones still run, so one unavailable
|
||||
# subsystem costs a section of the report rather than the whole thing.
|
||||
#
|
||||
# Flag Pass-Through
|
||||
# --dry-run and --log are forwarded to every child script.
|
||||
#
|
||||
# Notification Contract
|
||||
# notify() fires on failure only outside --dry-run, matching the runtime-mode contract
|
||||
# below — a dry run never sends anything outward.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# COFFEE_REPORT_SCRIPTS — ordered list of monitor scripts to run
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# sunday_morning_coffee_report.sh
|
||||
# Normal run.
|
||||
#
|
||||
# sunday_morning_coffee_report.sh --dry-run
|
||||
# Preview without any writes or notifications.
|
||||
#
|
||||
# sunday_morning_coffee_report.sh --log
|
||||
# Verbose per-script output.
|
||||
#
|
||||
# sunday_morning_coffee_report.sh --status
|
||||
# Show configured scripts and exit.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# sunday_morning_coffee_report.sh — normal run
|
||||
# sunday_morning_coffee_report.sh --dry-run — preview without any writes or notifications
|
||||
# sunday_morning_coffee_report.sh --log — verbose per-script output
|
||||
# sunday_morning_coffee_report.sh --status — show configured scripts and exit
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
source "$ECOSYSTEM_ROOT/load_config.sh"
|
||||
|
||||
SCRIPTS_ROOT="$SCRIPT_DIR/.."
|
||||
# Timed from here so the standard summary can report a real duration; this report had none.
|
||||
REPORT_START=$(date +%s)
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "$EUID" -ne 0 ]]; then
|
||||
error "Must be run as root"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#COFFEE_REPORT_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "COFFEE_REPORT_SCRIPTS is empty — no coffee report scripts will run"
|
||||
error "Check COFFEE_REPORT_SCRIPTS in master.conf"
|
||||
notify "coffee report scripts skipped on $(hostname) ($MY_ID) — COFFEE_REPORT_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Helpers ━━━
|
||||
# ==============================================================================================
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
|
||||
run_job() {
|
||||
local script_entry="$1"
|
||||
local extra_dry=""
|
||||
local extra_log=""
|
||||
[[ "$DRY_RUN" == true ]] && extra_dry="--dry-run"
|
||||
[[ "$ENABLE_LOGGING" == true ]] && extra_log="--log"
|
||||
|
||||
read -r -a script_args <<< "$script_entry"
|
||||
local script_path="$SCRIPTS_ROOT/${script_args[0]}"
|
||||
local script_name
|
||||
script_name=$(basename "${script_args[0]}")
|
||||
local extra_args=("${script_args[@]:1}")
|
||||
|
||||
if [[ ! -f "$script_path" ]]; then
|
||||
error "$script_name — not found at $script_path"
|
||||
JOB_FAIL+=("$script_name")
|
||||
return 1
|
||||
fi
|
||||
|
||||
log "Running: $script_name ${extra_args[*]}"
|
||||
if bash "$script_path" "${extra_args[@]}" $extra_dry $extra_log; then
|
||||
log "$script_name — done ✅"
|
||||
JOB_PASS+=("$script_name")
|
||||
else
|
||||
error "$script_name — failed (exit $?)"
|
||||
JOB_FAIL+=("$script_name")
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
@@ -87,11 +152,21 @@ fi
|
||||
# ==============================================================================================
|
||||
# ━━━ Main ━━━
|
||||
# ==============================================================================================
|
||||
_unraid_ver=$(platform_get_os_version 2>/dev/null || echo "unknown")
|
||||
_uptime_s=$(awk '{print int($1)}' /proc/uptime 2>/dev/null || echo 0)
|
||||
_container_count=$(docker ps -q 2>/dev/null | wc -l || echo 0)
|
||||
_rootfs_pct=$(df / --output=pcent 2>/dev/null | tail -1 | tr -d ' %')
|
||||
|
||||
echo ""
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
echo "☕ Sunday Morning Coffee Report — $MY_ID"
|
||||
[[ "$DRY_RUN" == true ]] && echo " [DRY RUN]"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
log "$ICON_HOST Server: $MY_ID ($LOCAL_SERVER_NAME) — unRAID $_unraid_ver"
|
||||
log "$ICON_TIME Uptime: $(format_duration $_uptime_s)"
|
||||
log "$ICON_CONTAINERS Docker: $_container_count container(s) running"
|
||||
log "$ICON_HEALTH Rootfs: ${_rootfs_pct:-?}% used"
|
||||
log "$ICON_GEAR Scripts: ${#COFFEE_REPORT_SCRIPTS[@]} configured"
|
||||
|
||||
if [[ ${#COFFEE_REPORT_SCRIPTS[@]} -eq 0 ]]; then
|
||||
warn "No scripts configured — add entries to COFFEE_REPORT_SCRIPTS in master.conf"
|
||||
@@ -101,7 +176,7 @@ fi
|
||||
for script_entry in "${COFFEE_REPORT_SCRIPTS[@]}"; do
|
||||
[[ -z "$script_entry" ]] && continue
|
||||
echo ""
|
||||
run_job "$script_entry"
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -112,6 +187,19 @@ echo "━━━━━ ☕ COFFEE REPORT SUMMARY ━━━━━"
|
||||
echo "🖥️ Host: $MY_ID"
|
||||
echo "⏱️ Done: $(date '+%Y-%m-%d %H:%M:%S')"
|
||||
|
||||
# Anything the Scheduler troubleshooter concluded was a Varaverk defect rather than a setting.
|
||||
# Counted here because a finding filed on a Tuesday and read on a Tuesday is a finding nobody
|
||||
# acts on — this report is the weekly moment the operator is actually looking. Count only: the
|
||||
# detail lives on the AI tab, and a report that reprints every open bug stops being skimmable.
|
||||
# Silent when there are none, and silent when AI is off, so a host without it reads the same as
|
||||
# it always has.
|
||||
if [[ "${AI_ENABLED:-false}" == true && -d "$DATA_DIR/ai_bugs" ]]; then
|
||||
_ai_bugs_open=$(grep -l '"open": true' "$DATA_DIR"/ai_bugs/*.json 2>/dev/null | wc -l)
|
||||
if [[ "${_ai_bugs_open:-0}" -gt 0 ]]; then
|
||||
echo "🐞 AI bugs: $_ai_bugs_open open — see the AI tab"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ${#JOB_PASS[@]} -gt 0 ]]; then
|
||||
echo "✅ Passed: ${JOB_PASS[*]}"
|
||||
fi
|
||||
@@ -119,5 +207,9 @@ if [[ ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
echo "❌ Failed: ${JOB_FAIL[*]}"
|
||||
fi
|
||||
|
||||
[[ ${#JOB_FAIL[@]} -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
# Standard ending. The configured section list is the denominator, so a report that quietly
|
||||
# stopped producing one of its sections reads as skipped rather than simply not appearing.
|
||||
JOB_COUNT="${#SUNDAY_REPORT_SCRIPTS[@]:-0}"
|
||||
[[ "$JOB_COUNT" -eq 0 ]] && JOB_COUNT=$(( ${#JOB_PASS[@]} + ${#JOB_FAIL[@]} ))
|
||||
orchestrator_summary "SUNDAY MORNING COFFEE REPORT" "$REPORT_START" "Sunday Morning Coffee Report"
|
||||
exit $?
|
||||
|
||||
Regular → Executable
+147
-71
@@ -2,66 +2,123 @@
|
||||
# ==============================================================================================
|
||||
# ============================= Transcode Management ===========================================
|
||||
# ==============================================================================================
|
||||
# Orchestrator — runs transcode_cleanup.sh then transcode_manager.sh in the correct order.
|
||||
# Replace individual transcode_manager and transcode_cleanup cron entries with this.
|
||||
# Schedule: */7 * * * * (every 7 minutes via User Scripts plugin)
|
||||
#
|
||||
# ── WHY CLEANUP BEFORE MANAGER ────────────────────────────────────────────────────────────────
|
||||
# Cleanup runs first — removes stale segment files from ended sessions.
|
||||
# Manager runs after — threshold decisions based on real current usage post-cleanup.
|
||||
# Without this order, stale files inflate the ramdisk usage reading and trigger
|
||||
# unnecessary SSD flips even when active sessions would fit on the ramdisk.
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Runs TRANSCODE_MANAGEMENT_SCRIPTS in order each cron cycle — this is the
|
||||
# single cron entry replacing individual entries for each child script.
|
||||
# Schedule: */7 * * * * (every 7 minutes)
|
||||
#
|
||||
# ── WHAT EACH SCRIPT DOES ─────────────────────────────────────────────────────────────────────
|
||||
# transcode_cleanup.sh — removes aged segment files not open by any process
|
||||
# uses lsof for O(1) per-file active check (never per-file lsof)
|
||||
# also triggers flip-back to ramdisk after cleanup if recovered ✅
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# transcode_manager.sh — checks ramdisk usage against thresholds
|
||||
# flips symlink between ramdisk and SSD as needed
|
||||
# writes one entry to TRANSCODE_DAILY_LOG after each run
|
||||
# shows active Emby sessions with play method
|
||||
# Order driven by TRANSCODE_MANAGEMENT_SCRIPTS in master.conf.
|
||||
# Default: transcode_cleanup.sh → transcode_manager.sh
|
||||
#
|
||||
# ── DAILY LOG ─────────────────────────────────────────────────────────────────────────────────
|
||||
# transcode_manager.sh writes to TRANSCODE_DAILY_LOG after each run:
|
||||
# transcode_cleanup.sh
|
||||
# Removes aged segment files not open by any process. Uses lsof for O(1)
|
||||
# per-file active check. Triggers flip-back to ramdisk after cleanup if
|
||||
# the ramdisk usage has recovered.
|
||||
#
|
||||
# transcode_manager.sh
|
||||
# Checks ramdisk usage against thresholds. Flips the symlink between ramdisk
|
||||
# and SSD as needed. Writes one entry to TRANSCODE_DAILY_LOG after each run.
|
||||
# Shows active Emby sessions with play method.
|
||||
#
|
||||
# DAILY LOG (written by transcode_manager.sh, not this orchestrator):
|
||||
# Format: DATE|RAMDISK_USED_GB|FLIP_COUNT|RAM_SESSIONS|SSD_SESSIONS
|
||||
# This orchestrator does NOT write its own log — manager handles it ✅
|
||||
# Log trimmed to TRANSCODE_LOG_RETENTION days by manager on each write.
|
||||
# Read by sunday_morning_coffee_report.sh and weekly_health_digest.sh.
|
||||
# Trimmed to TRANSCODE_LOG_RETENTION days on each write.
|
||||
# Read by sunday_morning_coffee_report.sh and weekly_health_digest.sh.
|
||||
#
|
||||
# ── HOST AWARENESS ────────────────────────────────────────────────────────────────────────────
|
||||
# detect_hosts() sets MY_ID and aliases RAMDISK_PATH, TRANSCODE_SSD, RAMDISK_WARN_GB etc.
|
||||
# Each server manages its own transcode location independently.
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── SAFEGUARDS ────────────────────────────────────────────────────────────────────────────────
|
||||
# Root check — mount and docker operations require root
|
||||
# acquire_lock — prevents concurrent 3-minute cycles overlapping
|
||||
# detect_hosts() — correct paths per host
|
||||
# --dry-run — passed through to both child scripts
|
||||
# Exit code — worst exit code of both scripts returned
|
||||
# Cleanup First, Decide Later
|
||||
# Stale segment files from ended sessions inflate the ramdisk usage reading
|
||||
# and trigger unnecessary SSD flips even when active sessions would fit on
|
||||
# the ramdisk. Cleanup runs first so the manager measures real current usage.
|
||||
# Order is config-driven (TRANSCODE_MANAGEMENT_SCRIPTS) but this dependency
|
||||
# is real — reordering the array changes what the manager measures.
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# TRANSCODE_DAILY_LOG — daily stats log (written by transcode_manager.sh)
|
||||
# Delegated Logging
|
||||
# This orchestrator does not write its own log — transcode_manager.sh owns
|
||||
# the TRANSCODE_DAILY_LOG write. One log writer, one format, no duplication.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Mount and docker operations require root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents concurrent 7-minute cycles overlapping. Cleanup and the manager
|
||||
# both touch the same ramdisk, and two cycles at once could have one deleting files
|
||||
# while the other is measuring usage to decide whether to flip.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() aliases RAMDISK_PATH, TRANSCODE_SSD and RAMDISK_WARN_GB per host.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if TRANSCODE_MANAGEMENT_SCRIPTS is empty —
|
||||
# without it the ramdisk would silently stop being cleaned or flipped, and the first
|
||||
# symptom would be a full ramdisk stalling playback.
|
||||
#
|
||||
# Ordering Is Load-Bearing
|
||||
# Cleanup runs before the manager so the manager measures real active-session usage
|
||||
# rather than usage inflated by stale files. Reversing them would trigger flips that
|
||||
# a cleanup two seconds later would have made unnecessary.
|
||||
#
|
||||
# Dry Run Propagation
|
||||
# --dry-run is passed through to every script in TRANSCODE_MANAGEMENT_SCRIPTS.
|
||||
#
|
||||
# Any-Failure Exit Code
|
||||
# Exits 1 if any child failed, 0 otherwise — the individual exit codes are not
|
||||
# propagated, only whether anything failed. A failure in an early child is therefore
|
||||
# never masked by a later success.
|
||||
#
|
||||
# Notification Contract
|
||||
# notify() fires on failure and is skipped in --dry-run.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# TRANSCODE_MANAGEMENT_SCRIPTS — scripts to run, in order
|
||||
# TRANSCODE_DAILY_LOG — daily stats log (written by transcode_manager.sh)
|
||||
# TRANSCODE_LOG_RETENTION — days to keep (trimmed by manager)
|
||||
# TRANSCODE_STATE_FILE — current state (ramdisk_used, flip_count etc.)
|
||||
# All TRANSCODE_* threshold vars — see master.conf Transcode Manager section
|
||||
# TRANSCODE_STATE_FILE — current state (ramdisk_used, flip_count, etc.)
|
||||
# TRANSCODE_* threshold vars — see master.conf Transcode Manager section
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# transcode_management.sh
|
||||
# Normal run (every 7 minutes via cron).
|
||||
#
|
||||
# transcode_management.sh --dry-run
|
||||
# Preview without changes (passed to every script in TRANSCODE_MANAGEMENT_SCRIPTS).
|
||||
#
|
||||
# transcode_management.sh --status
|
||||
# Show configuration and current state.
|
||||
#
|
||||
# transcode_management.sh --log
|
||||
# Verbose output from every child script.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# transcode_management.sh — normal run (every 7 minutes via cron)
|
||||
# transcode_management.sh --dry-run — preview without changes (passed to children)
|
||||
# transcode_management.sh --status — show configuration and current state
|
||||
# transcode_management.sh --log — verbose output from both child scripts
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
source "$ECOSYSTEM_ROOT/load_config.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
CLEANUP_SCRIPT="$SCRIPT_DIR/../Transcodes/transcode_cleanup.sh"
|
||||
MANAGER_SCRIPT="$SCRIPT_DIR/../Transcodes/transcode_manager.sh"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Setup ━━━
|
||||
# ==============================================================================================
|
||||
@@ -80,6 +137,20 @@ fi
|
||||
# detect_hosts() sets MY_ID and aliases all HOST*_TRANSCODE_* vars
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
# This orchestrator never timed itself, so its summary could not report a duration. Set before
|
||||
# any work so the figure means the cycle, not the tail of it.
|
||||
CYCLE_START=$(date +%s)
|
||||
|
||||
if [[ ${#TRANSCODE_MANAGEMENT_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "TRANSCODE_MANAGEMENT_SCRIPTS is empty — no transcode management scripts will run"
|
||||
error "Check TRANSCODE_MANAGEMENT_SCRIPTS in master.conf"
|
||||
notify "transcode management scripts skipped on $(hostname) ($MY_ID) — TRANSCODE_MANAGEMENT_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — passing through to child scripts"
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -95,12 +166,15 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
echo "$ICON_TIME Schedule: every 7 minutes"
|
||||
echo ""
|
||||
echo "━━━ Child Scripts ━━━"
|
||||
[[ -f "$CLEANUP_SCRIPT" ]] && \
|
||||
echo " $ICON_SUCCESS transcode_cleanup.sh — found" || \
|
||||
echo " $ICON_ERROR transcode_cleanup.sh — NOT FOUND at $CLEANUP_SCRIPT"
|
||||
[[ -f "$MANAGER_SCRIPT" ]] && \
|
||||
echo " $ICON_SUCCESS transcode_manager.sh — found" || \
|
||||
echo " $ICON_ERROR transcode_manager.sh — NOT FOUND at $MANAGER_SCRIPT"
|
||||
for _entry in "${TRANSCODE_MANAGEMENT_SCRIPTS[@]}"; do
|
||||
_script_path="$SCRIPT_DIR/../$_entry"
|
||||
_script_name=$(basename "$_entry")
|
||||
if [[ -f "$_script_path" ]]; then
|
||||
echo " $ICON_SUCCESS $_script_name — found"
|
||||
else
|
||||
echo " $ICON_ERROR $_script_name — NOT FOUND at $_script_path"
|
||||
fi
|
||||
done
|
||||
echo ""
|
||||
echo "━━━ Daily Log ━━━"
|
||||
if [[ -f "${TRANSCODE_DAILY_LOG:-}" ]] && [[ -s "$TRANSCODE_DAILY_LOG" ]]; then
|
||||
@@ -125,38 +199,40 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Validate Child Scripts ━━━
|
||||
# ━━━ Pre-run State Snapshot ━━━
|
||||
# ==============================================================================================
|
||||
if [[ ! -f "$CLEANUP_SCRIPT" ]]; then
|
||||
error "transcode_cleanup.sh not found: $CLEANUP_SCRIPT"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ! -f "$MANAGER_SCRIPT" ]]; then
|
||||
error "transcode_manager.sh not found: $MANAGER_SCRIPT"
|
||||
exit 1
|
||||
if mountpoint -q "${RAMDISK_PATH:-}" 2>/dev/null; then
|
||||
_rd_used=$(df "$RAMDISK_PATH" --output=used 2>/dev/null | tail -1 | tr -d ' ')
|
||||
_rd_avail=$(df "$RAMDISK_PATH" --output=avail 2>/dev/null | tail -1 | tr -d ' ')
|
||||
_rd_used_gb=$(awk "BEGIN {printf \"%.2f\", ${_rd_used:-0}/1048576}")
|
||||
_rd_avail_gb=$(awk "BEGIN {printf \"%.2f\", ${_rd_avail:-0}/1048576}")
|
||||
_rd_target=$(readlink "${TRANSCODE_LINK:-}" 2>/dev/null || echo "unknown")
|
||||
log "$ICON_RAM Ramdisk: ${_rd_used_gb}GB used / ${_rd_avail_gb}GB avail — symlink → ${_rd_target##*/}"
|
||||
else
|
||||
log "$ICON_RAM Ramdisk: not mounted"
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Run Cleanup ━━━
|
||||
# ━━━ Run Scripts ━━━
|
||||
# ==============================================================================================
|
||||
DRY_FLAG=""
|
||||
[[ "$DRY_RUN" == true ]] && DRY_FLAG="--dry-run"
|
||||
|
||||
bash "$CLEANUP_SCRIPT" $DRY_FLAG
|
||||
CLEANUP_EXIT=$?
|
||||
# transcode_manager.sh writes to TRANSCODE_DAILY_LOG after each run — no flag here
|
||||
# suppresses that; it owns the log write for this cycle regardless of position ✅
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
for _entry in "${TRANSCODE_MANAGEMENT_SCRIPTS[@]}"; do
|
||||
[[ -z "$_entry" ]] && continue
|
||||
run_orch_child "$_entry"
|
||||
done
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Run Manager ━━━
|
||||
# ━━━ Summary — minimal one-liner by default (7-min cadence — keep it quiet when healthy) ━━━
|
||||
# ==============================================================================================
|
||||
# transcode_manager.sh writes to TRANSCODE_DAILY_LOG after each run
|
||||
# No --no-log flag here — manager owns the log write for this cycle ✅
|
||||
bash "$MANAGER_SCRIPT" $DRY_FLAG
|
||||
MANAGER_EXIT=$?
|
||||
# Quiet by default — 7-min cadence. Anything failed or skipped expands on its own.
|
||||
JOB_COUNT="${#TRANSCODE_MANAGEMENT_SCRIPTS[@]}"
|
||||
orchestrator_summary "TRANSCODE CYCLE" "${CYCLE_START:-$(date +%s)}" "Transcode Management" quiet
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Exit ━━━
|
||||
# ==============================================================================================
|
||||
# Return worst exit code — caller knows if either script failed
|
||||
[[ "$CLEANUP_EXIT" -ne 0 || "$MANAGER_EXIT" -ne 0 ]] && exit 1
|
||||
[[ "${#JOB_FAIL[@]}" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
@@ -2,45 +2,109 @@
|
||||
# ==============================================================================================
|
||||
# ============================ Watchdog Orchestrator ===========================================
|
||||
# ==============================================================================================
|
||||
# Runs WATCHDOG_ORCHESTRATOR_SCRIPTS in order each cron cycle.
|
||||
# Schedule: * * * * * (every minute via User Scripts plugin)
|
||||
#
|
||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||
# Driven by WATCHDOG_ORCHESTRATOR_SCRIPTS in master.conf — add, remove, or reorder there.
|
||||
# Default: resource_watchdog → docker_watchdog → system_watchdog
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Runs WATCHDOG_ORCHESTRATOR_SCRIPTS in order each cron cycle. Replaces the
|
||||
# continuous loops previously embedded in individual watchdog scripts — those
|
||||
# are now single-pass; this orchestrator provides the cadence.
|
||||
# Schedule: */15 * * * * (every 15 minutes)
|
||||
#
|
||||
# ── WHY ORDER MATTERS ─────────────────────────────────────────────────────────────────────────
|
||||
# Resource Watchdog first — frees RAM and CPU before healing attempts container restarts.
|
||||
# Containers restarted into a resource-pressured system just fail again.
|
||||
# Docker Watchdog second — restarts with pressure already reduced, more likely to stabilise.
|
||||
# System Watchdog last — only triggers if prior layers could not resolve the issue.
|
||||
# Rebooting without first reducing pressure may reboot into the same state.
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── STARTUP GRACE ─────────────────────────────────────────────────────────────────────────────
|
||||
# No action until system uptime >= WATCHDOG_STARTUP_GRACE seconds.
|
||||
# Prevents false positives from containers still starting at array launch.
|
||||
# Each sub-script enforces this independently — orchestrator exits early to avoid log noise.
|
||||
# Order driven by WATCHDOG_ORCHESTRATOR_SCRIPTS in master.conf.
|
||||
# Default: resource_watchdog → docker_watchdog → system_watchdog →
|
||||
# unraid_api_key_renew → stability_watchdog
|
||||
#
|
||||
# ── OVERLAP PROTECTION ────────────────────────────────────────────────────────────────────────
|
||||
# acquire_lock() — exits immediately if a prior cycle is still in progress.
|
||||
# Prevents pile-up when a cycle runs long (daemon restart attempt = 30s, etc.).
|
||||
# ARRAY CHECK
|
||||
# Exits immediately if /mnt/user is not mounted as shfs. Watchdogs check
|
||||
# Docker containers and storage — meaningless without the array. Prevents
|
||||
# false positives and unnecessary reboots when array is stopped or stopping.
|
||||
#
|
||||
# ── REPLACES ──────────────────────────────────────────────────────────────────────────────────
|
||||
# Continuous loops previously in system_watchdog.sh and docker_watchdog.sh.
|
||||
# Those scripts are now single-pass — this orchestrator provides the cadence.
|
||||
# Remove system_watchdog.sh and docker_watchdog.sh from ARRAY_START_SCRIPTS.
|
||||
# STARTUP GRACE
|
||||
# No action until system uptime >= WATCHDOG_STARTUP_GRACE seconds. Prevents
|
||||
# false positives from containers still starting at array launch. Each
|
||||
# sub-script enforces this independently — orchestrator exits early to avoid
|
||||
# log noise.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Pressure Before Healing
|
||||
# Resource Watchdog runs first — it frees RAM and CPU before any container
|
||||
# restart is attempted. Containers restarted into a resource-pressured system
|
||||
# just fail again. Docker Watchdog restarts with pressure already reduced.
|
||||
# System Watchdog checks component health after containers are healed.
|
||||
# Stability Watchdog reboots only when all prior layers could not resolve the
|
||||
# issue. API key renew is check-first and silent when valid.
|
||||
#
|
||||
# Never Queue
|
||||
# acquire_lock exits immediately if a prior cycle is still running. Prevents
|
||||
# pile-up when a cycle runs long (daemon restart attempt = 30s, etc.).
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Every watchdog launched from here requires root. Failing once at the top gives one
|
||||
# clear error instead of the same permission failure repeated per child.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock in strict mode — if the previous cycle is still running, this one exits
|
||||
# rather than queuing. At a 15-minute cadence a waiting lock would pile up cycles
|
||||
# behind a slow watchdog and eventually run them all at once.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID for logs and notifications.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if WATCHDOG_ORCHESTRATOR_SCRIPTS is empty.
|
||||
# Without it the cycle reports "0/0 passed" and exits 0 every cycle — indistinguishable
|
||||
# from a healthy run, while nothing at all is being monitored.
|
||||
#
|
||||
# Array Check
|
||||
# Exits early if /mnt/user is not shfs-mounted. Watchdogs that inspect shares would
|
||||
# otherwise read an unmounted array as missing data and act on it.
|
||||
#
|
||||
# Startup Grace
|
||||
# WATCHDOG_STARTUP_GRACE is respected before any checks run, so containers still
|
||||
# initialising after boot are not judged as unhealthy.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# run_orch_child() records a failing or missing watchdog and continues. One broken
|
||||
# watchdog never suppresses the rest of the chain.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# WATCHDOG_ORCHESTRATOR_SCRIPTS — watchdogs to run, in order
|
||||
# WATCHDOG_STARTUP_GRACE — seconds after boot before checks activate
|
||||
# WATCHDOG_ORCHESTRATOR_HEARTBEAT — periodic heartbeat log toggle
|
||||
# WATCHDOG_ORCHESTRATOR_HEARTBEAT_HOURS — heartbeat interval in hours
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# watchdog_orchestrator.sh — normal run (called by cron every minute)
|
||||
# watchdog_orchestrator.sh --dry-run — pass --dry-run to all sub-scripts
|
||||
# watchdog_orchestrator.sh --status — show script paths and current grace state
|
||||
# watchdog_orchestrator.sh --log — verbose output from all sub-scripts
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# watchdog_orchestrator.sh
|
||||
# Normal run (called by cron every 15 minutes).
|
||||
#
|
||||
# watchdog_orchestrator.sh --dry-run
|
||||
# Pass --dry-run to all sub-scripts.
|
||||
#
|
||||
# watchdog_orchestrator.sh --status
|
||||
# Show script paths and current grace state.
|
||||
#
|
||||
# watchdog_orchestrator.sh --log
|
||||
# Verbose output from all sub-scripts.
|
||||
#
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
@@ -63,6 +127,19 @@ acquire_lock
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#WATCHDOG_ORCHESTRATOR_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "WATCHDOG_ORCHESTRATOR_SCRIPTS is empty — no watchdogs will run"
|
||||
error "Check WATCHDOG_ORCHESTRATOR_SCRIPTS in master.conf"
|
||||
notify "watchdogs skipped on $(hostname) ($MY_ID) — WATCHDOG_ORCHESTRATOR_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
log "$ICON_GEAR Config: grace=${WATCHDOG_STARTUP_GRACE}s heartbeat=${WATCHDOG_ORCHESTRATOR_HEARTBEAT:-true}/${WATCHDOG_ORCHESTRATOR_HEARTBEAT_HOURS:-1}hr scripts=${#WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"
|
||||
log "$ICON_WATCHDOG Order: $(for s in "${WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"; do printf '%s ' "${s##*/}"; done)"
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — passing --dry-run to all sub-scripts"
|
||||
|
||||
# Derive a display name from a script path: "resource_watchdog.sh" → "Resource Watchdog"
|
||||
@@ -104,56 +181,41 @@ if [[ "$SHOW_STATUS" == true ]]; then
|
||||
done
|
||||
|
||||
echo ""
|
||||
echo " Schedule: * * * * * (every minute via User Scripts)"
|
||||
echo " Schedule: */15 * * * * (every 15 minutes)"
|
||||
echo " Heartbeat: ${WATCHDOG_ORCHESTRATOR_HEARTBEAT:-true} / every ${WATCHDOG_ORCHESTRATOR_HEARTBEAT_HOURS:-1}hr"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Array Check ━━━
|
||||
# ==============================================================================================
|
||||
if ! platform_storage_healthy; then
|
||||
echo "Array not started — skipping watchdog cycle"
|
||||
exit 0
|
||||
fi
|
||||
log "$ICON_DISK Array: $(platform_storage_path) mounted ✅"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Startup Grace ━━━
|
||||
# ==============================================================================================
|
||||
UPTIME_SECONDS=$(awk '{print int($1)}' /proc/uptime)
|
||||
if [[ "$UPTIME_SECONDS" -lt "$WATCHDOG_STARTUP_GRACE" ]]; then
|
||||
log "Startup grace — ${UPTIME_SECONDS}s / ${WATCHDOG_STARTUP_GRACE}s — skipping cycle"
|
||||
echo "Startup grace — $(format_duration $UPTIME_SECONDS) / $(format_duration $WATCHDOG_STARTUP_GRACE) — skipping cycle"
|
||||
exit 0
|
||||
fi
|
||||
log "Startup grace: past — uptime $(format_duration $UPTIME_SECONDS)"
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Run Watchdog Cycle ━━━
|
||||
# ==============================================================================================
|
||||
CYCLE_START=$(date +%s)
|
||||
PASS=()
|
||||
FAIL=()
|
||||
|
||||
run_watchdog() {
|
||||
local name="$1" script="$2"
|
||||
|
||||
if [[ ! -f "$script" ]]; then
|
||||
error "$name — not found: $script"
|
||||
FAIL+=("$name:missing")
|
||||
return 1
|
||||
fi
|
||||
|
||||
[[ ! -x "$script" ]] && chmod +x "$script"
|
||||
|
||||
local extra_args=()
|
||||
[[ "$DRY_RUN" == true ]] && extra_args+=("--dry-run")
|
||||
[[ "$VERBOSE" == true ]] && extra_args+=("--log")
|
||||
|
||||
log "$ICON_START $name"
|
||||
if bash "$script" "${extra_args[@]}"; then
|
||||
PASS+=("$name")
|
||||
return 0
|
||||
else
|
||||
error "$name — non-zero exit"
|
||||
FAIL+=("$name")
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
JOB_PASS=()
|
||||
JOB_FAIL=()
|
||||
|
||||
for _entry in "${WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"; do
|
||||
run_watchdog "$(_watchdog_display_name "$_entry")" "$ECOSYSTEM_ROOT/$_entry"
|
||||
[[ -z "$_entry" ]] && continue
|
||||
run_orch_child "$_entry"
|
||||
done
|
||||
|
||||
CYCLE_END=$(date +%s)
|
||||
@@ -164,7 +226,7 @@ DURATION=$(( CYCLE_END - CYCLE_START ))
|
||||
# ==============================================================================================
|
||||
if [[ "${WATCHDOG_ORCHESTRATOR_HEARTBEAT:-true}" == true ]]; then
|
||||
HB_SECONDS=$(( ${WATCHDOG_ORCHESTRATOR_HEARTBEAT_HOURS:-1} * 3600 ))
|
||||
HB_COUNT_FILE="/tmp/watchdog_orch_hb.count"
|
||||
HB_COUNT_FILE="${STATE_DIR:-/tmp}/watchdog_orch_hb.count"
|
||||
HB_COUNT=$(cat "$HB_COUNT_FILE" 2>/dev/null || echo 0)
|
||||
HB_COUNT=$(( HB_COUNT + 1 ))
|
||||
echo "$HB_COUNT" > "$HB_COUNT_FILE"
|
||||
@@ -172,23 +234,22 @@ if [[ "${WATCHDOG_ORCHESTRATOR_HEARTBEAT:-true}" == true ]]; then
|
||||
HB_ELAPSED=$(( HB_COUNT * 60 ))
|
||||
if [[ "$HB_SECONDS" -gt 0 ]] && (( HB_ELAPSED % HB_SECONDS < 60 )) && [[ "$HB_COUNT" -gt 1 ]]; then
|
||||
HB_HR=$(( HB_ELAPSED / 3600 ))
|
||||
warn "♥ watchdog_orchestrator alive — $MY_ID — ~${HB_HR}hr ($(date '+%H:%M:%S'))"
|
||||
log "♥ watchdog_orchestrator alive — $MY_ID — ~${HB_HR}hr ($(date '+%H:%M:%S'))"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary — only shown on failures or --log ━━━
|
||||
# ━━━ Summary — minimal one-liner by default, full breakdown on failure or --log ━━━
|
||||
# ==============================================================================================
|
||||
if [[ "${#FAIL[@]}" -gt 0 || "$VERBOSE" == true ]]; then
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY WATCHDOG CYCLE — $MY_ID — $(date '+%H:%M:%S') ━━━━━"
|
||||
for p in "${PASS[@]}"; do log " $ICON_DONE $p"; done
|
||||
for f in "${FAIL[@]}"; do error " $ICON_ERROR $f"; done
|
||||
echo "$ICON_TIME Duration: $(format_duration $DURATION)"
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
if [[ "${#FAIL[@]}" -gt 0 ]]; then
|
||||
notify "Watchdog cycle failure on $(hostname) ($MY_ID) — ${FAIL[*]}" \
|
||||
"Watchdog Orchestrator" "warning"
|
||||
fi
|
||||
# Per-script detail only when there is something to read; the standard block carries the rest.
|
||||
if [[ "${#JOB_FAIL[@]}" -gt 0 || "$ENABLE_LOGGING" == true ]]; then
|
||||
for p in "${JOB_PASS[@]}"; do log " $ICON_DONE $p"; done
|
||||
for f in "${JOB_FAIL[@]}"; do error " $ICON_ERROR $f"; done
|
||||
fi
|
||||
|
||||
# Quiet by default at a 15-min cadence. The configured script list is the denominator, so a
|
||||
# watchdog that silently stopped running one of its checks shows up as skipped.
|
||||
JOB_COUNT="${#WATCHDOG_ORCHESTRATOR_SCRIPTS[@]}"
|
||||
_mode=quiet; [[ "$ENABLE_LOGGING" == true ]] && _mode=full
|
||||
orchestrator_summary "WATCHDOG CYCLE" "$CYCLE_START" "Watchdog Orchestrator" "$_mode"
|
||||
exit $?
|
||||
|
||||
Regular → Executable
+152
-132
@@ -2,66 +2,110 @@
|
||||
# ==============================================================================================
|
||||
# ============================= Weekly Sync Maintenance ========================================
|
||||
# ==============================================================================================
|
||||
# Weekly maintenance window orchestrator — clean sync, container updates, weekly restarts.
|
||||
# Schedule: 30 2 * * 0 (Sunday 2:30am — before Sunday 7am coffee report)
|
||||
#
|
||||
# ── EXECUTION ORDER ───────────────────────────────────────────────────────────────────────────
|
||||
# 1. Stop local containers — Emby + auth stack stopped locally
|
||||
# 2. Stop remote containers — Emby + auth stack stopped remotely via SSH
|
||||
# 3. Pull updates locally — if WEEKLY_SYNC_UPDATES=true (zero extra downtime)
|
||||
# 4. Pull updates remotely — if WEEKLY_SYNC_UPDATES_REMOTE=true
|
||||
# 5. rsync WEEKLY_SYNC_SHARES — full clean mirror, containers stopped both sides
|
||||
# 6. Start remote containers — correct order, delayed start respected
|
||||
# 7. Start local containers — correct order, delayed start respected
|
||||
# 8. WEEKLY_MAINTENANCE_SCRIPTS — weekly restarts etc. (docker_weekly_restart.sh)
|
||||
# 9. docker_update.sh --remainder — update all containers not in daily or weekly sync window
|
||||
# PURPOSE
|
||||
# ─────────────────────────────────────────────────────────────────────────────
|
||||
# Weekly maintenance window orchestrator — clean Emby sync, container updates,
|
||||
# and weekly restarts. Schedule: 30 2 * * 0 (Sunday 2:30am)
|
||||
#
|
||||
# ── WHY WEEKLY NOT NIGHTLY FOR EMBY ──────────────────────────────────────────────────────────
|
||||
# Emby builds a warm image cache on HOST2 throughout the week.
|
||||
# Syncing nightly resets cache — cold loads every morning for users.
|
||||
# Weekly sync: cache stays warm 6 days, resets Sunday night while users sleep.
|
||||
# emby-fallback dirty sync covers watch states + library every 30min between weekly syncs.
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL MODEL
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── CONTAINER UPDATES ─────────────────────────────────────────────────────────────────────────
|
||||
# Containers already stopped for sync — updates pull at zero extra downtime.
|
||||
# Both servers start on identical image versions after the window completes.
|
||||
# Toggle: WEEKLY_SYNC_UPDATES / WEEKLY_SYNC_UPDATES_REMOTE in master.conf
|
||||
# 1. Stop local containers — Emby + auth stack stopped locally
|
||||
# 2. Stop remote containers — Emby + auth stack stopped remotely via SSH
|
||||
# 3. Pull updates locally — if WEEKLY_SYNC_UPDATES=true (zero extra downtime)
|
||||
# 4. Pull updates remotely — if WEEKLY_SYNC_UPDATES_REMOTE=true
|
||||
# 5. rsync WEEKLY_SYNC_SHARES — full clean mirror, containers stopped both sides
|
||||
# 6. Start remote containers — correct order, delayed start respected
|
||||
# 7. Start local containers — rebuild if new image pulled, docker start otherwise
|
||||
# 8. WEEKLY_MAINTENANCE_SCRIPTS — weekly restarts etc. (docker_weekly_restart.sh)
|
||||
#
|
||||
# ── HOST AWARENESS ────────────────────────────────────────────────────────────────────────────
|
||||
# detect_hosts() sets MY_ID — used in banner, summary, and notifications.
|
||||
# WEEKLY_SYNC_SHARES and WEEKLY_MAINTENANCE_SCRIPTS configured in master.conf.
|
||||
# Same script runs correctly on both servers.
|
||||
# ==============================================================================================
|
||||
# DESIGN PRINCIPLES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# ── SAFEGUARDS ────────────────────────────────────────────────────────────────────────────────
|
||||
# Root check — stop/start containers, rsync require root
|
||||
# acquire_lock — prevents concurrent weekly windows
|
||||
# check_connectivity — verifies remote before any remote operations
|
||||
# check_remote_rootfs — aborts if remote rootfs nearly full
|
||||
# DOCKER_TIMEOUT — all docker calls protected
|
||||
# SSH_TIMEOUT — all SSH calls protected
|
||||
# validate_unraid_cmd — notify validated before use
|
||||
# Silent on success — runs weekly, only failures warrant notification
|
||||
# Weekly Cadence Preserves Cache
|
||||
# Emby builds a warm image cache on HOST2 throughout the week. Syncing nightly
|
||||
# resets that cache — cold loads every morning for users. Weekly sync keeps
|
||||
# the cache warm for 6 days, resets Sunday night while users sleep.
|
||||
# play_state_sync covers watch/resume state every 30 min between weekly syncs.
|
||||
#
|
||||
# ── CONFIGURATION (master.conf) ───────────────────────────────────────────────────────────────
|
||||
# WEEKLY_SYNC_SHARES — shares synced during window
|
||||
# WEEKLY_MAINTENANCE_SCRIPTS — scripts run after sync
|
||||
# WEEKLY_SYNC_UPDATES — toggle local container updates
|
||||
# WEEKLY_SYNC_UPDATES_REMOTE — toggle remote container updates
|
||||
# WEEKLY_RSYNC_ENABLED — enable/disable rsync section
|
||||
# Zero-Downtime Updates
|
||||
# Containers are already stopped for the sync window — image pulls happen at
|
||||
# zero extra downtime. Both servers start on identical image versions after
|
||||
# the window completes.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# OPERATIONAL SAFEGUARDS
|
||||
# ==============================================================================================
|
||||
#
|
||||
# Root Enforcement
|
||||
# Container stop/start and rsync both require root.
|
||||
#
|
||||
# Lock Acquisition
|
||||
# acquire_lock prevents concurrent weekly windows. This window stops Emby and the auth
|
||||
# stack — two overlapping runs would fight over the same critical containers.
|
||||
#
|
||||
# Host Detection
|
||||
# detect_hosts() sets MY_ID for the banner, summary and notifications.
|
||||
#
|
||||
# Empty Job List Guard
|
||||
# Exits with an error and a notification if WEEKLY_MAINTENANCE_SCRIPTS is empty, rather
|
||||
# than taking the weekly outage window and doing nothing with it.
|
||||
#
|
||||
# Connectivity Check
|
||||
# check_connectivity verifies the remote before any remote operation is attempted.
|
||||
#
|
||||
# Remote Rootfs Check
|
||||
# check_remote_rootfs aborts rsync if the remote rootfs is nearly full.
|
||||
#
|
||||
# Timeout Protection
|
||||
# DOCKER_TIMEOUT bounds every docker call and SSH_TIMEOUT every SSH call, so neither a
|
||||
# hung daemon nor an unresponsive partner can hold the weekly window open indefinitely.
|
||||
#
|
||||
# Non-Fatal Steps
|
||||
# A failing job is recorded and the remaining jobs still run.
|
||||
#
|
||||
# Silent on Success
|
||||
# Runs weekly; only failures warrant a notification.
|
||||
#
|
||||
# ==============================================================================================
|
||||
# CONFIGURATION
|
||||
# ==============================================================================================
|
||||
#
|
||||
# master.conf
|
||||
#
|
||||
# WEEKLY_SYNC_SHARES — shares synced during window
|
||||
# WEEKLY_MAINTENANCE_SCRIPTS — scripts run after sync
|
||||
# WEEKLY_SYNC_UPDATES — toggle local container updates
|
||||
# WEEKLY_SYNC_UPDATES_REMOTE — toggle remote container updates
|
||||
# WEEKLY_RSYNC_ENABLED — enable/disable rsync section
|
||||
#
|
||||
# ==============================================================================================
|
||||
# RUNTIME MODES
|
||||
# ==============================================================================================
|
||||
#
|
||||
# weekly_sync_maintenance.sh
|
||||
# Normal run.
|
||||
#
|
||||
# weekly_sync_maintenance.sh --dry-run
|
||||
# Preview without stopping containers or syncing.
|
||||
#
|
||||
# weekly_sync_maintenance.sh --log
|
||||
# Verbose per-share/per-job output.
|
||||
#
|
||||
# weekly_sync_maintenance.sh --status
|
||||
# Show configuration and exit.
|
||||
#
|
||||
# ── USAGE ─────────────────────────────────────────────────────────────────────────────────────
|
||||
# weekly_sync_maintenance.sh — normal run
|
||||
# weekly_sync_maintenance.sh --dry-run — preview without stopping containers or syncing
|
||||
# weekly_sync_maintenance.sh --log — verbose per-share/per-job output
|
||||
# weekly_sync_maintenance.sh --status — show configuration and exit
|
||||
# ==============================================================================================
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
|
||||
source "$SCRIPT_DIR/../load_config.sh"
|
||||
|
||||
RSYNC_SCRIPT="$SCRIPT_DIR/../Rsync/rsync.sh"
|
||||
SCRIPTS_ROOT="$SCRIPT_DIR/.."
|
||||
ECOSYSTEM_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
|
||||
RSYNC_SCRIPT="$ECOSYSTEM_ROOT/Rsync/rsync.sh"
|
||||
|
||||
parse_args "$@"
|
||||
|
||||
@@ -76,10 +120,6 @@ if [[ "$EUID" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
validate_unraid_cmd \
|
||||
"/usr/local/emhttp/plugins/dynamix/scripts/notify" \
|
||||
"" "" \
|
||||
"unRAID notify script" || warn "unRAID notify script not found — native notifications disabled"
|
||||
|
||||
acquire_lock
|
||||
|
||||
@@ -89,6 +129,16 @@ if ! command -v docker &>/dev/null; then
|
||||
fi
|
||||
|
||||
detect_hosts
|
||||
|
||||
# An unconfigured job list would run nothing and still report "0/0 passed" — indistinguishable
|
||||
# from a healthy run. Fail loudly instead of silently doing no work.
|
||||
if [[ ${#WEEKLY_MAINTENANCE_SCRIPTS[@]} -eq 0 ]]; then
|
||||
error "WEEKLY_MAINTENANCE_SCRIPTS is empty — no weekly maintenance scripts will run"
|
||||
error "Check WEEKLY_MAINTENANCE_SCRIPTS in master.conf"
|
||||
notify "weekly maintenance scripts skipped on $(hostname) ($MY_ID) — WEEKLY_MAINTENANCE_SCRIPTS is empty" \
|
||||
"$(basename "$0" .sh)" "warning"
|
||||
exit 1
|
||||
fi
|
||||
resolve_remote_ip
|
||||
|
||||
WINDOW_START=$(date +%s)
|
||||
@@ -103,35 +153,6 @@ read -r -a MAINTENANCE_CONTAINERS <<< \
|
||||
|
||||
[[ "$DRY_RUN" == true ]] && warn "DRY RUN — no containers will be stopped, no sync, no updates"
|
||||
|
||||
# ── Helper — run a post-sync maintenance script ────────────────────────────────────────────────
|
||||
run_job() {
|
||||
local script_entry="$1"
|
||||
local extra_dry=""
|
||||
[[ "$DRY_RUN" == true ]] && extra_dry="--dry-run"
|
||||
|
||||
read -r -a script_args <<< "$script_entry"
|
||||
local script_path="$SCRIPTS_ROOT/${script_args[0]}"
|
||||
local script_name
|
||||
script_name=$(basename "${script_args[0]}")
|
||||
local extra_args=("${script_args[@]:1}")
|
||||
|
||||
if [[ ! -f "$script_path" ]]; then
|
||||
error "$script_name — not found at $script_path"
|
||||
JOB_FAIL+=("$script_name")
|
||||
return 1
|
||||
fi
|
||||
|
||||
log "Running: $script_name ${extra_args[*]}"
|
||||
# shellcheck disable=SC2086
|
||||
if bash "$script_path" "${extra_args[@]}" $extra_dry; then
|
||||
log "$script_name — done ✅"
|
||||
JOB_PASS+=("$script_name ${extra_args[*]}")
|
||||
else
|
||||
error "$script_name — failed (exit $?)"
|
||||
JOB_FAIL+=("$script_name ${extra_args[*]}")
|
||||
fi
|
||||
}
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Status ━━━
|
||||
# ==============================================================================================
|
||||
@@ -194,6 +215,8 @@ else
|
||||
# Load container lists for stop functions
|
||||
read -r -a CRITICAL_CONTAINER_NAMES <<< \
|
||||
"${PROFILE_CRITICAL_CONTAINER_NAMES[critical-data]:-} ${PROFILE_CRITICAL_CONTAINER_NAMES[emby]:-}"
|
||||
read -r -a LOCAL_CRITICAL_CONTAINER_NAMES <<< \
|
||||
"${PROFILE_CRITICAL_CONTAINER_NAMES[critical-data]:-} ${PROFILE_CRITICAL_CONTAINER_NAMES[emby]:-}"
|
||||
read -r -a DELAYED_CONTAINERS <<< "${PROFILE_DELAYED_CONTAINERS[critical-data]:-}"
|
||||
CONTAINER_DELAY="${PROFILE_CONTAINER_DELAY[critical-data]:-15}"
|
||||
|
||||
@@ -207,6 +230,8 @@ fi
|
||||
echo ""
|
||||
echo "━━━ $ICON_GEAR Container Updates ━━━"
|
||||
|
||||
declare -A _weekly_needs_rebuild=()
|
||||
|
||||
if [[ "$WEEKLY_SYNC_UPDATES" == true ]]; then
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
for c in "${MAINTENANCE_CONTAINERS[@]}"; do
|
||||
@@ -222,13 +247,21 @@ if [[ "$WEEKLY_SYNC_UPDATES" == true ]]; then
|
||||
log "$c — not found locally, skipping update"
|
||||
continue
|
||||
fi
|
||||
_old_id=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
log "Pulling $IMAGE for $c..."
|
||||
if docker pull "$IMAGE" >/dev/null 2>&1; then
|
||||
log "$c — image updated ✅"
|
||||
_new_id=$(docker image inspect "$IMAGE" --format='{{.Id}}' 2>/dev/null || echo "")
|
||||
if [[ -n "$_old_id" && "$_old_id" != "$_new_id" ]]; then
|
||||
log "$c — new image (${_old_id:7:12} → ${_new_id:7:12}) — will rebuild after sync"
|
||||
_weekly_needs_rebuild["$c"]=1
|
||||
else
|
||||
log "$c — already current"
|
||||
fi
|
||||
else
|
||||
warn "$c — pull failed, will start on existing image"
|
||||
fi
|
||||
done
|
||||
unset _old_id _new_id
|
||||
fi
|
||||
else
|
||||
echo "WEEKLY_SYNC_UPDATES=false — skipping local updates"
|
||||
@@ -254,7 +287,7 @@ if [[ "$WEEKLY_SYNC_UPDATES_REMOTE" == true ]]; then
|
||||
-o ConnectTimeout="$SSH_TIMEOUT" \
|
||||
root@"$REMOTE_SERVER" \
|
||||
"docker pull $IMAGE" >/dev/null 2>&1; then
|
||||
log "$c — remote image updated ✅"
|
||||
echo "$c — remote image updated ✅"
|
||||
else
|
||||
warn "$c — remote pull failed, will start on existing image"
|
||||
fi
|
||||
@@ -299,7 +332,7 @@ else
|
||||
|
||||
if [[ "$EXIT_CODE" -eq 0 ]]; then
|
||||
PASS+=("$JOB_NAME")
|
||||
log "$JOB_NAME — done in $JOB_DUR ✅"
|
||||
echo "$JOB_NAME — done in $JOB_DUR ✅"
|
||||
else
|
||||
FAIL+=("$JOB_NAME")
|
||||
error "$JOB_NAME — failed after $JOB_DUR (exit $EXIT_CODE)"
|
||||
@@ -320,7 +353,35 @@ if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — containers will not be started"
|
||||
else
|
||||
start_containers
|
||||
start_local_containers
|
||||
|
||||
# Local start — rebuild containers that received a new image, docker start the rest
|
||||
if [[ ${#LOCAL_RUNNING_CONTAINERS[@]} -eq 0 ]]; then
|
||||
log "No local containers to restart."
|
||||
else
|
||||
for _c in "${LOCAL_RUNNING_CONTAINERS[@]}"; do
|
||||
[[ -z "$_c" ]] && continue
|
||||
_needs_delay=false
|
||||
for _d in "${DELAYED_CONTAINERS[@]}"; do
|
||||
[[ "$_c" == "$_d" ]] && _needs_delay=true && break
|
||||
done
|
||||
[[ "$_needs_delay" == true ]] && {
|
||||
info "Waiting ${CONTAINER_DELAY}s before starting $_c..."
|
||||
sleep "$CONTAINER_DELAY"
|
||||
}
|
||||
if [[ -n "${_weekly_needs_rebuild[$_c]:-}" ]]; then
|
||||
log "Rebuilding $_c on new image..."
|
||||
if platform_rebuild_container "$_c"; then
|
||||
echo "$_c rebuilt on new image ✅"
|
||||
else
|
||||
warn "$_c rebuild failed — falling back to docker start"
|
||||
docker start "$_c" >/dev/null 2>&1 || error "Failed to start $_c"
|
||||
fi
|
||||
else
|
||||
docker start "$_c" >/dev/null 2>&1 && echo "$_c started" || error "Failed to start $_c"
|
||||
fi
|
||||
done
|
||||
unset _c _d _needs_delay
|
||||
fi
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
@@ -332,52 +393,23 @@ if [[ ${#WEEKLY_MAINTENANCE_SCRIPTS[@]} -gt 0 ]]; then
|
||||
for script_entry in "${WEEKLY_MAINTENANCE_SCRIPTS[@]}"; do
|
||||
[[ -z "$script_entry" ]] && continue
|
||||
echo ""
|
||||
run_job "$script_entry"
|
||||
run_orch_child "$script_entry"
|
||||
done
|
||||
fi
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Remainder Container Updates ━━━
|
||||
# ==============================================================================================
|
||||
# Updates all running containers not already covered by daily or the weekly sync window.
|
||||
# Runs last — weekly sync-window containers are back up before this pulls their peers.
|
||||
echo ""
|
||||
echo "━━━ $ICON_CONTAINERS Remainder Container Updates ━━━"
|
||||
|
||||
DOCKER_UPDATE_SCRIPT="$SCRIPTS_ROOT/Docker_Essentials/docker_update.sh"
|
||||
if [[ ! -f "$DOCKER_UPDATE_SCRIPT" ]]; then
|
||||
warn "docker_update.sh not found — skipping remainder updates"
|
||||
JOB_FAIL+=("docker_update.sh --remainder")
|
||||
else
|
||||
_remainder_args=("--remainder")
|
||||
[[ "$DRY_RUN" == true ]] && _remainder_args+=("--dry-run")
|
||||
if bash "$DOCKER_UPDATE_SCRIPT" "${_remainder_args[@]}"; then
|
||||
echo "Remainder updates complete ✅"
|
||||
JOB_PASS+=("docker_update.sh --remainder")
|
||||
else
|
||||
warn "Remainder updates completed with errors"
|
||||
JOB_FAIL+=("docker_update.sh --remainder")
|
||||
fi
|
||||
unset _remainder_args
|
||||
fi
|
||||
|
||||
WINDOW_END=$(date +%s)
|
||||
|
||||
# ==============================================================================================
|
||||
# ━━━ Summary ━━━
|
||||
# ==============================================================================================
|
||||
# Per-unit detail first — the standard block that follows carries the verdict and the counts, not
|
||||
# the names, and knowing WHICH share failed is the whole point of reading a log.
|
||||
echo ""
|
||||
echo "━━━━━ $ICON_SUMMARY WEEKLY SYNC MAINTENANCE SUMMARY ━━━━━"
|
||||
echo "$ICON_HOST Identity: $MY_ID ($LOCAL_SERVER_NAME)"
|
||||
echo "$ICON_TIME Window: $(date -d @"$WINDOW_START" '+%Y-%m-%d %H:%M:%S') → $(date -d @"$WINDOW_END" '+%H:%M:%S')"
|
||||
echo "$ICON_TIME Duration: $(format_duration $(( WINDOW_END - WINDOW_START )))"
|
||||
echo "$ICON_GEAR Updates: local=${WEEKLY_SYNC_UPDATES:-false} remote=${WEEKLY_SYNC_UPDATES_REMOTE:-false}"
|
||||
echo ""
|
||||
|
||||
echo "$ICON_SYNC Sync jobs ($SHARE_COUNT):"
|
||||
for job in "${PASS[@]}"; do echo " $ICON_DONE $job"; done
|
||||
for job in "${FAIL[@]}"; do echo " $ICON_ERROR $job"; done
|
||||
echo " Passed: ${#PASS[@]} Failed: ${#FAIL[@]}"
|
||||
|
||||
if [[ ${#JOB_PASS[@]} -gt 0 || ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
echo ""
|
||||
@@ -386,19 +418,7 @@ if [[ ${#JOB_PASS[@]} -gt 0 || ${#JOB_FAIL[@]} -gt 0 ]]; then
|
||||
for job in "${JOB_FAIL[@]}"; do echo " $ICON_ERROR $job"; done
|
||||
fi
|
||||
|
||||
TOTAL_FAIL=$(( ${#FAIL[@]} + ${#JOB_FAIL[@]} ))
|
||||
|
||||
echo ""
|
||||
if [[ "$DRY_RUN" == true ]]; then
|
||||
warn "DRY RUN — no changes made"
|
||||
elif [[ "$TOTAL_FAIL" -eq 0 ]]; then
|
||||
echo "$ICON_DONE Status: all complete ✅ — ${#PASS[@]} share(s) synced, ${#JOB_PASS[@]} job(s) run"
|
||||
else
|
||||
warn "Status: $TOTAL_FAIL failure(s)"
|
||||
notify "Weekly maintenance failed on $(hostname) ($MY_ID) — sync: ${#FAIL[@]}/$SHARE_COUNT failed, jobs: ${#JOB_FAIL[@]} failed" \
|
||||
"Weekly Maintenance" "warning"
|
||||
fi
|
||||
echo "━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━"
|
||||
|
||||
[[ "$TOTAL_FAIL" -gt 0 ]] && exit 1
|
||||
exit 0
|
||||
# Standard ending. Derives skipped from SHARE_COUNT vs what actually ran, so a gated-off section
|
||||
# can no longer read as success — this is the run that printed "all complete — 0 shares synced".
|
||||
orchestrator_summary "WEEKLY SYNC MAINTENANCE" "$WINDOW_START" "Weekly Maintenance"
|
||||
exit $?
|
||||
@@ -138,7 +138,7 @@ PARTNERSHIP_ONBOARD_NOTIFY=true # notify both servers on successful onboard
|
||||
# Format: "ContainerName|WebUIPort"
|
||||
HOST1_PARTNERSHIP_AUTH_WEBUIS=(
|
||||
"NginxProxyManager|81"
|
||||
"Lldap-Gmer4Lfe|17170"
|
||||
"Lldap|17170"
|
||||
"Authelia|9091"
|
||||
"Authelia-Secondary|9092"
|
||||
)
|
||||
@@ -154,7 +154,7 @@ HOST1_PARTNERSHIP_AUTH_STACK=(
|
||||
"my-Authelia.xml"
|
||||
"my-Authelia-Secondary.xml"
|
||||
"my-NginxProxyManager.xml"
|
||||
"my-Lldap-Gmer4Lfe.xml"
|
||||
"my-Lldap.xml"
|
||||
)
|
||||
|
||||
# XML templates pushed to mirror during onboard (arr stack).
|
||||
@@ -430,6 +430,29 @@ Called automatically by `partnership_manager.sh --offboard`. Can also be run dir
|
||||
| `--status` | Show state files, blocklist, SSH key status |
|
||||
| `--unblock <hostname>` | Remove hostname from blocklist |
|
||||
|
||||
### partnership_transfer.sh
|
||||
|
||||
Owner-only. Transfers ownership to the current mirror — swaps roles without moving
|
||||
containers. Run via `partnership_manager.sh --transfer --confirm=<phrase>`.
|
||||
|
||||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--confirm=<phrase>` | Required confirmation string (from `PARTNERSHIP_TRANSFER_CONFIRM` in master.conf) |
|
||||
| `--dry-run` | Preview all steps without executing |
|
||||
| `--log` | Verbose per-step output |
|
||||
|
||||
### onboard_cancel.sh
|
||||
|
||||
Removes SSH keys between hosts in the specified direction. Safe to run at any onboard
|
||||
phase — clears the corresponding setup.db flags.
|
||||
|
||||
| Flag | Effect |
|
||||
|------|--------|
|
||||
| `--direction=h1` | Remove HOST1 → HOST2 key (default) |
|
||||
| `--direction=h2` | Remove HOST2 → HOST1 key |
|
||||
| `--direction=both` | Both directions |
|
||||
| `--dry-run` | Preview without making changes |
|
||||
|
||||
### ssh_setup.sh
|
||||
|
||||
| Flag | Effect |
|
||||
@@ -551,7 +574,7 @@ Partnership/partnership_manager.sh --status # shows both sides via SSH
|
||||
|
||||
```bash
|
||||
# Check the offline counter:
|
||||
cat /boot/config/partnership_offline_days.db
|
||||
cat "$STATE_DIR/partnership_offline_days.db"
|
||||
|
||||
# Extended Tailscale outage may have incremented the counter.
|
||||
# Check Tailscale peer visibility:
|
||||
@@ -568,9 +591,9 @@ Partnership/partnership_manager.sh --onboard
|
||||
## ━━━ STATE FILES ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
```bash
|
||||
/boot/config/partnership_HOST1.db # HOST1 writes only
|
||||
/boot/config/partnership_HOST2.db # HOST2 writes only
|
||||
/boot/config/partnership_blocklist.db # hostname|timestamp|reason
|
||||
$STATE_DIR/partnership_HOST1.db # HOST1 writes only
|
||||
$STATE_DIR/partnership_HOST2.db # HOST2 writes only
|
||||
$STATE_DIR/partnership_blocklist.db # hostname|timestamp|reason
|
||||
|
||||
# Example state file:
|
||||
state=ACTIVE
|
||||
|
||||
@@ -141,7 +141,11 @@ independently at 4-hour cadence.
|
||||
| `partnership_onboard.sh` | One-time setup — SSH keys, stack deploy, arr bootstrap | Manually, once per server per partnership |
|
||||
| `partnership_offboard.sh` | Clean separation — both paths, both roles | Via `partnership_manager.sh --offboard`; or directly |
|
||||
| `partnership_manager.sh` | Dispatcher + monitor — onboard WebUIs, health check, transfer, status | `--check` every 30min; all other modes manually |
|
||||
| `partnership_transfer.sh` | Transfer ownership from current owner to current mirror | Via `partnership_manager.sh --transfer`; owner only |
|
||||
| `ssh_setup.sh` | SSH key generation, remote install, auth validation | Called by onboard; manually for re-keying or validation |
|
||||
| `onboard_cancel.sh` | Remove SSH keys in one or both directions, clear setup flags | During cancelled or failed onboard; manual cleanup |
|
||||
| `gitea_ssh_setup.sh` | Generate a Gitea keypair and register it via the Gitea API | Called during onboard; manually when re-keying a server |
|
||||
| `share_setup.sh` | Create missing Unraid shares on the mirror from the owner's sync lists | Called during onboard; safe to re-run — existing shares are never modified |
|
||||
|
||||
---
|
||||
|
||||
@@ -157,7 +161,6 @@ HOST2 (mirror) runs:
|
||||
HOST1 (owner) runs:
|
||||
partnership_onboard.sh
|
||||
├─ ssh_setup.sh generates keypair, copies to mirror
|
||||
├─ [plugin install on mirror] FolderView3 if configured
|
||||
├─ [stop mirror auth stack] PARTNERSHIP_REPLACE_CONTAINERS via SSH
|
||||
├─ deploy_container_from_xml() pushes auth XMLs to mirror + starts containers
|
||||
│ └─ wait_for_container_healthy() Mariadb/Redis health-checked before Authelia
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user