Record measured Ollama tuning results in the AI design notes
The estimates in the original note were wrong in both directions, and Unraid's Apply does not reliably recreate a container, which is worth knowing before the next tuning pass.
This commit is contained in:
+78
-21
@@ -1,7 +1,11 @@
|
|||||||
# Varaverk AI Integration — Design Notes
|
# Varaverk AI Integration — Design Notes
|
||||||
|
|
||||||
**Status: design only. Nothing below is built.** Ollama is installed and running on HOST1;
|
**Status: design only. Nothing below is built.** No Varaverk script calls Ollama, and no
|
||||||
no Varaverk script calls it. Captured 2026-08-01 so the reasoning survives.
|
`AI_*` variable exists in any conf yet. Captured 2026-08-01 so the reasoning survives.
|
||||||
|
|
||||||
|
Ollama itself *is* installed, tuned and verified on HOST1 — `qwen2.5-coder:14b` for
|
||||||
|
generation, `nomic-embed-text` for embeddings, 16k context, pinned to the RTX 3080. See
|
||||||
|
Hardware Budget for measured numbers. That is the substrate, not the integration.
|
||||||
|
|
||||||
**Origin:** the RTX 3080 was freed when the Windows gaming VM was retired. It is bound to the
|
**Origin:** the RTX 3080 was freed when the Windows gaming VM was retired. It is bound to the
|
||||||
`nvidia` driver, not `vfio` — not reserved for passthrough, so there is no VM contention to
|
`nvidia` driver, not `vfio` — not reserved for passthrough, so there is no VM contention to
|
||||||
@@ -85,7 +89,11 @@ it is already waiting on.
|
|||||||
# ━━━ AI ━━━
|
# ━━━ AI ━━━
|
||||||
AI_ENABLED=false # master switch — fail-closed, != "true" means off
|
AI_ENABLED=false # master switch — fail-closed, != "true" means off
|
||||||
AI_CONNECT_TIMEOUT=5 # probe timeout when resolving a host
|
AI_CONNECT_TIMEOUT=5 # probe timeout when resolving a host
|
||||||
AI_REQUEST_TIMEOUT=120 # generation can be slow; cold load was 28.6s
|
AI_REQUEST_TIMEOUT=240 # must clear a cold load — measured 1m45s after tuning.
|
||||||
|
# KEEP_ALIVE=-1 means this only bites after a restart,
|
||||||
|
# but a first call that times out is the worst first
|
||||||
|
# impression a caller can have. Re-measure against a
|
||||||
|
# real RAG query before fixing this number.
|
||||||
AI_RESOLVE_CACHE_TTL=300 # don't re-probe the mesh on every script invocation
|
AI_RESOLVE_CACHE_TTL=300 # don't re-probe the mesh on every script invocation
|
||||||
AI_MAX_RETRIES=1 # AI is enhancement — do not retry hard
|
AI_MAX_RETRIES=1 # AI is enhancement — do not retry hard
|
||||||
|
|
||||||
@@ -229,8 +237,11 @@ strictly wider than required, for no benefit. Restrict to loopback + Tailscale,
|
|||||||
Tailscale ACLs to allow only mesh nodes. Node-level ACLs fit the mesh model better than
|
Tailscale ACLs to allow only mesh nodes. Node-level ACLs fit the mesh model better than
|
||||||
app-level auth Ollama cannot provide anyway.
|
app-level auth Ollama cannot provide anyway.
|
||||||
|
|
||||||
Also set: `/ext-varaverk` mount is now `ro` in the template (was `rw` into live prod).
|
**Done 2026-08-01:** `/ext-varaverk` is now mounted `ro` (was `rw` into live prod).
|
||||||
**Requires an Apply in the Docker tab to take effect on the running container.**
|
Verified on the running container — `rw=false`.
|
||||||
|
|
||||||
|
**Still open:** the LAN exposure above. Deliberately not folded into the tuning rebuild,
|
||||||
|
since bind-address versus Tailscale ACL is a decision rather than a setting.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -239,29 +250,75 @@ Also set: `/ext-varaverk` mount is now `ro` in the template (was `rw` into live
|
|||||||
RTX 3080, 10 GB, pinned to Ollama by UUID — isolated from the Quadro P2000 that Emby
|
RTX 3080, 10 GB, pinned to Ollama by UUID — isolated from the Quadro P2000 that Emby
|
||||||
transcodes on. Do not let AI onto the P2000.
|
transcodes on. Do not let AI onto the P2000.
|
||||||
|
|
||||||
| | VRAM |
|
**Tuned and measured 2026-08-01.** These are observed values, not estimates.
|
||||||
|---|---|
|
|
||||||
| `qwen2.5-coder:14b` (Q4_K_M) | 8.34 GB |
|
|
||||||
| `nomic-embed-text` (768-dim) | 0.25 GB |
|
|
||||||
| **used / total** | **9.30 / 10.24 GB (91%)** |
|
|
||||||
|
|
||||||
Latency: 28.6s cold load, ~2.6s warm, ~4.4s warm with both models resident.
|
| | before | after |
|
||||||
`OLLAMA_KEEP_ALIVE=-1` keeps both pinned, so cold load is a boot-time cost only.
|
|---|---|---|
|
||||||
|
| `OLLAMA_NUM_PARALLEL` | 2 | **1** |
|
||||||
|
| `OLLAMA_KV_CACHE_TYPE` | f16 | **q8_0** |
|
||||||
|
| `OLLAMA_CONTEXT_LENGTH` | 4096 | **16384** |
|
||||||
|
| `OLLAMA_FLASH_ATTENTION` | false | **true** |
|
||||||
|
| VRAM used | 9298 MiB (91%) | **8811 MiB (86%)** |
|
||||||
|
| warm latency | 2.6s | **1.85s** |
|
||||||
|
| cold load | 28.6s | **1m45s** |
|
||||||
|
|
||||||
|
`/api/ps` confirms `ctx=16384` — the increase is real, not just an env var.
|
||||||
|
|
||||||
|
**4× the context for less VRAM than before.** Flash Attention plus the quantized KV cache
|
||||||
|
more than paid for the increase. Cold load got much slower, which is irrelevant while
|
||||||
|
`OLLAMA_KEEP_ALIVE=-1` pins both models — but it is felt after any container restart.
|
||||||
|
|
||||||
|
### Flash Attention is mandatory, not optional
|
||||||
|
|
||||||
|
`OLLAMA_KV_CACHE_TYPE=q8_0` **will not load** without it:
|
||||||
|
|
||||||
|
```
|
||||||
|
llama_init_from_model: V cache quantization requires flash_attn
|
||||||
|
llama-server process no longer running: exit status 1
|
||||||
|
```
|
||||||
|
|
||||||
|
Quantized K/V cache requires Flash Attention. Supported on Ampere and newer; the 3080
|
||||||
|
qualifies. If the KV cache type is ever changed back toward a quantized value, Flash
|
||||||
|
Attention must be on or the model silently fails to load and every call errors.
|
||||||
|
|
||||||
|
### Tuning order still matters
|
||||||
|
|
||||||
|
If these are ever re-tuned from defaults, the order is load-bearing — raising context first
|
||||||
|
at high utilisation will OOM:
|
||||||
|
|
||||||
|
1. `OLLAMA_NUM_PARALLEL` → 1 (each slot multiplies KV cache)
|
||||||
|
2. `OLLAMA_FLASH_ATTENTION` → true (prerequisite for the next step)
|
||||||
|
3. `OLLAMA_KV_CACHE_TYPE` → q8_0 (roughly halves KV memory)
|
||||||
|
4. *then* `OLLAMA_CONTEXT_LENGTH` upward
|
||||||
|
|
||||||
|
32k was considered and rejected — projected ~1000 MiB of KV, leaving under 450 MiB headroom.
|
||||||
|
16k is the comfortable ceiling for a 14B on this card.
|
||||||
|
|
||||||
|
### Concurrency
|
||||||
|
|
||||||
**You cannot have 14B + long context + real parallelism on 10 GB.** Pick two.
|
**You cannot have 14B + long context + real parallelism on 10 GB.** Pick two.
|
||||||
`OLLAMA_MAX_QUEUE=512` means excess requests queue rather than fail, and at ~4s responses,
|
`OLLAMA_MAX_QUEUE=512` means excess requests queue rather than fail, and at ~2s responses,
|
||||||
two or three users serialised is barely noticeable. Multi-user hits are expected to be rare.
|
two or three users serialised is barely noticeable. Multi-user hits are expected to be rare.
|
||||||
So: keep parallelism low, let the queue absorb bursts, spend VRAM on **context** — that is
|
So: parallelism stays at 1, the queue absorbs bursts, and the VRAM goes to **context** —
|
||||||
what RAG actually needs.
|
which is what RAG actually needs.
|
||||||
|
|
||||||
Tuning order matters. At 91% utilisation, raising context first will OOM:
|
### Applying template changes
|
||||||
|
|
||||||
1. `OLLAMA_NUM_PARALLEL` 2 → 1 (frees KV cache)
|
**Unraid's "Apply" does not reliably recreate the container.** Observed 2026-08-01: the
|
||||||
2. `OLLAMA_KV_CACHE_TYPE` f16 → q8_0 (roughly halves KV memory, negligible quality cost)
|
template was saved correctly but the container was only *restarted*, so the env vars never
|
||||||
3. *then* `OLLAMA_CONTEXT_LENGTH` 4096 → 8192/16384
|
took effect — `Created` stayed unchanged while `Started` advanced. Env changes require a
|
||||||
|
remove-and-recreate.
|
||||||
|
|
||||||
All three are env vars in the same template, so they ride along with the same Apply as the
|
Force it with Unraid's own script:
|
||||||
`ro` mount fix — one container recreate covers everything.
|
|
||||||
|
```bash
|
||||||
|
/usr/local/emhttp/plugins/dynamix.docker.manager/scripts/rebuild_container Ollama
|
||||||
|
docker start Ollama # rebuild stops it — Ollama is not in unraid-autostart
|
||||||
|
```
|
||||||
|
|
||||||
|
Verify with `docker inspect Ollama --format '{{.Created}}'` — the timestamp must move.
|
||||||
|
Checking the env vars alone is not enough; a restart leaves the old ones in place and
|
||||||
|
looks like nothing happened.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user