Record the RAG corpus shape while the audit context is fresh

The header audit and per-folder docs pass produced a corpus with properties
worth capturing before they are forgotten: deterministic chunk boundaries,
section type as a retrieval filter, and an index that cannot leak credentials
because the conf files were never tracked.
This commit is contained in:
Gmer4Lfe
2026-08-01 23:56:34 -04:00
parent c54039392d
commit ce806ae854
+103 -8
View File
@@ -332,20 +332,115 @@ No Python on Unraid, and none needed. Everything required is already present:
container needed
- **embed + generate** → Ollama HTTP, same `curl` pattern as every other integration here
Corpus: **49,654 lines of bash + ~11,970 lines of markdown.** Small.
**Chunking is already solved.** Every script now carries
`PURPOSE / OPERATIONAL MODEL / DESIGN PRINCIPLES / OPERATIONAL SAFEGUARDS / CONFIGURATION /
RUNTIME MODES` at exact, greppable boundaries. Those are semantically coherent units with
stable headings — far better retrieval chunks than fixed-size windows. The `DESIGN PRINCIPLES`
sections encode *why*, which is what the model needs and what code alone never says.
Note `qwen2.5-coder` returns `501 — does not support embeddings`. Embedding is
`nomic-embed-text`'s job. Batch embedding works (n inputs → n vectors in one call) and is
required — indexing 50k lines one HTTP call at a time is not viable.
---
### The corpus — and why its shape matters more than its size
As of 2026-08-01, after the header audit and the per-folder documentation pass:
| Layer | Size | What it answers |
|-------|------|-----------------|
| Script headers | 115 files × 6 sections = **690 chunks**, 14,685 lines | "What does *this script* do, and why that way" |
| Folder docs | 18 `README-*.md` + 13 `Manual-*.md` | "How does this *group* work" / "how do I do the thing" |
| Top-level | `README.md`, `Manual.md` | "What is this system" |
| Conf templates | 2,175 lines, **~55% comment** | The schema, self-describing |
| Bash bodies | 51,166 lines | Implementation — index last, lowest weight |
Markdown total: **13,516 lines.** Still small enough that cosine over the whole set is
milliseconds.
**Chunking is already solved, and the audit is what solved it.** Every script carries
`PURPOSE / OPERATIONAL MODEL / DESIGN PRINCIPLES / OPERATIONAL SAFEGUARDS / CONFIGURATION /
RUNTIME MODES`**115 of 115, no exceptions.** Split on `^# SECTION NAME$` and every chunk is
a semantically coherent unit by construction. The single worst failure mode in naive RAG —
a fixed-size window cutting mid-thought and embedding two half-ideas as one vector — cannot
happen here. Median header is 117 lines, so a section lands around 130200 tokens: comfortably
inside `nomic-embed-text`'s window, no sub-splitting needed.
**Store the section name as a column, not just as chunk text.** This is the highest-value
thing the audit bought and it should not be thrown away at index time. Section type is a free
metadata filter, so retrieval can route before it computes similarity:
| Question shape | Filter to |
|----------------|-----------|
| "what stops X and Y overlapping" | `OPERATIONAL SAFEGUARDS` |
| "what variable controls X" | `CONFIGURATION` |
| "does this take --dry-run" | `RUNTIME MODES` |
| "why is it built this way" | `DESIGN PRINCIPLES` |
| "what does this script do" | `PURPOSE` |
Hybrid retrieval essentially for free, because every chunk already has a type.
Suggested table shape:
```sql
CREATE TABLE vv_chunks (
id INTEGER PRIMARY KEY,
path TEXT NOT NULL, -- repo-relative
kind TEXT NOT NULL, -- header | readme | manual | template | body
section TEXT, -- PURPOSE, OPERATIONAL SAFEGUARDS, ... (NULL for md/body)
heading TEXT, -- md ## heading, for doc chunks
content TEXT NOT NULL,
vector BLOB NOT NULL, -- 768 float32
indexed INTEGER NOT NULL -- epoch; re-embed on mtime change only
);
```
### Why this corpus is worth more than an equivalent pile of code
A model can read `mover_stop.sh` and describe what it does. What it *cannot* derive from any
amount of source is that a thing was done deliberately. The audit wrote those down:
- the API cache writers are lockless and unprivileged **on purpose** — regenerable within a
minute, every consumer has a live fallback
- `removeCompletedDownloads` / `removeFailedDownloads` both true is **intended**, not an
oversight
- the arr cleanup ctime gate depends on `media_shares_permissions.sh` staying conditional —
reverting either silently stops orphan collection
- `mesh_monitor.sh`, `adapter.sh`, `decision_engine.sh`, `containers.sh` and
`api_cache_writer.sh` carry no root check and no lock **by design** — each documents why in
its own header (libraries that must not `exit`, read-only probes, or regenerable output
with a live fallback)
Without those in the index, the most likely contribution from an AI assistant reviewing this
repo is a confident regression: *"I notice this script lacks a lock."* Weight
`DESIGN PRINCIPLES` and `OPERATIONAL SAFEGUARDS` heavily for any suggest-a-change flow —
they are the guardrails against the assistant helpfully undoing a decision.
### Indexing is safe by default — keep it that way
`Configurations/*.conf` is gitignored; `Deployment/*.template` is tracked and carries all the
explanatory comments. The corpus therefore describes the full schema while structurally
**never containing a credential**, because the credential-bearing files were never in the repo
to begin with.
Treat that as a deliberate boundary, not a happy accident:
- **index tracked files only** — never walk `Configurations/`, `State_Files/`, or `data/`
- a live conf value that the model genuinely needs should arrive through a *tool call* at
query time, subject to the same redaction rules as everything else in the Security section,
not be baked into a vector at index time
- an embedded secret is unrevocable in a way a logged one is not — there is no rotation story
for a value already averaged into a 768-dim float
### Known gap — the PHP layer is not covered
78 PHP files under `Plugin/unraid/`; **2** carry a `PURPOSE` block. The entire web UI —
`pages/`, `api/`, `include/` — is effectively invisible to retrieval.
Consequence: any "AI helper per Varaverk page" feature has this as a hard prerequisite. A
page-scoped assistant that cannot retrieve the page's own logic is worse than no assistant.
`include/` is the high-value subset to do first — 16 files, and both the pages and the API
endpoints route through the same `vv_*()` builders, so documenting it once covers both
callers. This is a follow-on pass, not a blocker for indexing bash.
---
## Scheduled AI
Same orchestrator tiers as everything else, gated on `AI_ENABLED` plus a reachable host.