diff --git a/Notes_AI-Design.md b/Notes_AI-Design.md index 08cfb92..834eae3 100644 --- a/Notes_AI-Design.md +++ b/Notes_AI-Design.md @@ -332,20 +332,115 @@ No Python on Unraid, and none needed. Everything required is already present: container needed - **embed + generate** → Ollama HTTP, same `curl` pattern as every other integration here -Corpus: **49,654 lines of bash + ~11,970 lines of markdown.** Small. - -**Chunking is already solved.** Every script now carries -`PURPOSE / OPERATIONAL MODEL / DESIGN PRINCIPLES / OPERATIONAL SAFEGUARDS / CONFIGURATION / -RUNTIME MODES` at exact, greppable boundaries. Those are semantically coherent units with -stable headings — far better retrieval chunks than fixed-size windows. The `DESIGN PRINCIPLES` -sections encode *why*, which is what the model needs and what code alone never says. - Note `qwen2.5-coder` returns `501 — does not support embeddings`. Embedding is `nomic-embed-text`'s job. Batch embedding works (n inputs → n vectors in one call) and is required — indexing 50k lines one HTTP call at a time is not viable. --- +### The corpus — and why its shape matters more than its size + +As of 2026-08-01, after the header audit and the per-folder documentation pass: + +| Layer | Size | What it answers | +|-------|------|-----------------| +| Script headers | 115 files × 6 sections = **690 chunks**, 14,685 lines | "What does *this script* do, and why that way" | +| Folder docs | 18 `README-*.md` + 13 `Manual-*.md` | "How does this *group* work" / "how do I do the thing" | +| Top-level | `README.md`, `Manual.md` | "What is this system" | +| Conf templates | 2,175 lines, **~55% comment** | The schema, self-describing | +| Bash bodies | 51,166 lines | Implementation — index last, lowest weight | + +Markdown total: **13,516 lines.** Still small enough that cosine over the whole set is +milliseconds. + +**Chunking is already solved, and the audit is what solved it.** Every script carries +`PURPOSE / OPERATIONAL MODEL / DESIGN PRINCIPLES / OPERATIONAL SAFEGUARDS / CONFIGURATION / +RUNTIME MODES` — **115 of 115, no exceptions.** Split on `^# SECTION NAME$` and every chunk is +a semantically coherent unit by construction. The single worst failure mode in naive RAG — +a fixed-size window cutting mid-thought and embedding two half-ideas as one vector — cannot +happen here. Median header is 117 lines, so a section lands around 130–200 tokens: comfortably +inside `nomic-embed-text`'s window, no sub-splitting needed. + +**Store the section name as a column, not just as chunk text.** This is the highest-value +thing the audit bought and it should not be thrown away at index time. Section type is a free +metadata filter, so retrieval can route before it computes similarity: + +| Question shape | Filter to | +|----------------|-----------| +| "what stops X and Y overlapping" | `OPERATIONAL SAFEGUARDS` | +| "what variable controls X" | `CONFIGURATION` | +| "does this take --dry-run" | `RUNTIME MODES` | +| "why is it built this way" | `DESIGN PRINCIPLES` | +| "what does this script do" | `PURPOSE` | + +Hybrid retrieval essentially for free, because every chunk already has a type. + +Suggested table shape: + +```sql +CREATE TABLE vv_chunks ( + id INTEGER PRIMARY KEY, + path TEXT NOT NULL, -- repo-relative + kind TEXT NOT NULL, -- header | readme | manual | template | body + section TEXT, -- PURPOSE, OPERATIONAL SAFEGUARDS, ... (NULL for md/body) + heading TEXT, -- md ## heading, for doc chunks + content TEXT NOT NULL, + vector BLOB NOT NULL, -- 768 float32 + indexed INTEGER NOT NULL -- epoch; re-embed on mtime change only +); +``` + +### Why this corpus is worth more than an equivalent pile of code + +A model can read `mover_stop.sh` and describe what it does. What it *cannot* derive from any +amount of source is that a thing was done deliberately. The audit wrote those down: + +- the API cache writers are lockless and unprivileged **on purpose** — regenerable within a + minute, every consumer has a live fallback +- `removeCompletedDownloads` / `removeFailedDownloads` both true is **intended**, not an + oversight +- the arr cleanup ctime gate depends on `media_shares_permissions.sh` staying conditional — + reverting either silently stops orphan collection +- `mesh_monitor.sh`, `adapter.sh`, `decision_engine.sh`, `containers.sh` and + `api_cache_writer.sh` carry no root check and no lock **by design** — each documents why in + its own header (libraries that must not `exit`, read-only probes, or regenerable output + with a live fallback) + +Without those in the index, the most likely contribution from an AI assistant reviewing this +repo is a confident regression: *"I notice this script lacks a lock."* Weight +`DESIGN PRINCIPLES` and `OPERATIONAL SAFEGUARDS` heavily for any suggest-a-change flow — +they are the guardrails against the assistant helpfully undoing a decision. + +### Indexing is safe by default — keep it that way + +`Configurations/*.conf` is gitignored; `Deployment/*.template` is tracked and carries all the +explanatory comments. The corpus therefore describes the full schema while structurally +**never containing a credential**, because the credential-bearing files were never in the repo +to begin with. + +Treat that as a deliberate boundary, not a happy accident: + +- **index tracked files only** — never walk `Configurations/`, `State_Files/`, or `data/` +- a live conf value that the model genuinely needs should arrive through a *tool call* at + query time, subject to the same redaction rules as everything else in the Security section, + not be baked into a vector at index time +- an embedded secret is unrevocable in a way a logged one is not — there is no rotation story + for a value already averaged into a 768-dim float + +### Known gap — the PHP layer is not covered + +78 PHP files under `Plugin/unraid/`; **2** carry a `PURPOSE` block. The entire web UI — +`pages/`, `api/`, `include/` — is effectively invisible to retrieval. + +Consequence: any "AI helper per Varaverk page" feature has this as a hard prerequisite. A +page-scoped assistant that cannot retrieve the page's own logic is worse than no assistant. + +`include/` is the high-value subset to do first — 16 files, and both the pages and the API +endpoints route through the same `vv_*()` builders, so documenting it once covers both +callers. This is a follow-on pass, not a blocker for indexing bash. + +--- + ## Scheduled AI Same orchestrator tiers as everything else, gated on `AI_ENABLED` plus a reachable host.