[Sync] replace the flat layout with the unified stow tree

Supersedes the old flat .config/ layout (last published 2026-06-28) with the
private repo's structure: one shared base plus per-host overlays.

- packages: common/ gui/ lw/ fl/ wm/ plus install.sh and bin/ tooling
  (dotsync, reconcile-hyde.sh)
- new README covering the layout, deploy order and the HyDE dependency
- current HyDE waybar rig (layouts/, cava), pi agent extensions, claude/
  config, tmux, presenterm, aichat roles
- fish: kp (keepassxc-cli + fzf picker, db path from $KP_DB) and
  bind_M_n_history (alt+1..9 recalls the nth history entry)
- drops cruft that should never have been tracked: the duplicate top-level
  .pi/ copy, btop.log, zellij config.kdl.bak, fish_variables
- .pi/agent/auth.json is gitignored; auth.json.example ships instead

Host-specific work sessions and the personal backlog stay in the private
tree. Endpoint locators in the llamacpp/whisper guides are placeholders
($SERVER, <own-domain>) — the guides themselves stay, since they are the
useful part.
This commit is contained in:
coja
2026-08-13 02:53:06 +02:00
parent f62cb40499
commit 1e39e0cb4b
314 changed files with 12306 additions and 10772 deletions
+166
View File
@@ -0,0 +1,166 @@
# Local LLM server — setup & usage
llama.cpp in **router mode** serving ~10 model presets from one 16 GB GPU.
Last full tune: **2026-07-16** (decode numbers measured then via `bench.py`, unless a preset comment in `config.ini` says otherwise).
Last full clean sweep: **2026-08-08 03:22** — all 11 presets loaded and ran green (baseline 2.0 GB) after the week's coder-race/KV-fix/gemma-sidecar round.
- **Box**: AMD RX 7600 XT 16 GB (ROCm, ~288 GB/s) + Ryzen 5600X (6 cores) + 48 GB RAM.
⚠ The GPU **also drives the display** — see [VRAM safety](#vram-safety-the-golden-rules).
- **Endpoints**: LAN `http://$SERVER:11343/v1` · remote `https://<own-domain>/api/v1`
(own reverse proxy to the same server; requires an API key. Only the LAN endpoint is keyless).
`$SERVER` = the box's LAN address.
- **Files here**: `config.ini` (model presets — the section names ARE the API model ids),
`bench.py` (safety-first benchmark), `bench-results.md` / `bench-history.md` (ledgers, generated).
## How it runs
systemd unit **`llama.service`** (on the server, under a dedicated service user) runs a script wrapping:
```
llama-server --port 11343 --host 0.0.0.0 --models-max 1 --models-preset ~/.config/llamacpp/config.ini
```
- Router mode: each requested model loads on demand in a child process.
- **`--models-max 1` is deliberate and load-bearing**: at most one model resident; requesting
another LRU-evicts the current one *first*. Anything ≥2 lets two ~13 GB models stack →
VRAM overcommit → GTT spill → **whole-PC freeze**. Do not raise it casually.
(`sleep-idle-seconds` per preset is the second line of defense.)
- Models live in `~/software/models/` (paths in `config.ini` must stay absolute — llama-server
does not expand `~` or env vars in preset values). mmproj/draft files must match the exact
filename in the preset (watch for trailing spaces when renaming downloads!).
**Deploying config changes**: this dir is stowed into `~/.config/llamacpp/` as per-file
symlinks, so on the server `dotsync` (pull; re-stows only if files were added/removed) then
`sudo systemctl restart llama.service` — the router reads presets once at startup.
After ANY preset change: restart + `./bench.py -m <changed-ids>` and check the guards.
## VRAM safety (the golden rules)
ROCm does not OOM cleanly — an overcommitted model spills into GTT (system RAM),
starves the desktop and freezes the PC (reboot). Hard-won rules:
1. Keep every preset's benched **VRAM free ≥ 2.5 GB** (idle desktop uses ~1.3 of 17.2).
2. File size > VRAM is **fine** for MoE — `n-cpu-moe N` keeps the first N layers' experts
in system RAM. It is the main tuning knob: higher = safer/slower, lower = faster.
3. Bench with a **clean baseline** (close the browser): a dirty baseline both skews the
margins and costs ~34 t/s of decode (GPU compute contention — measured 2026-07-08).
4. Bench after every change; config-edit without restart = numbers from the *old* config.
## Model roster (reference numbers = clean night sweeps, 2026-07-16)
| model id | decode t/s | ctx | extras | role |
|---|---|---|---|---|
| `gemma-4-E4B-it-UD-Q8_K_XL` | **57.3** | **96k** | vision, MTP, prefill 185 t/s | fast small generalist, long docs, images (ctx 96k ✓ 08-08, n_ctx_train 131072) |
| `Qwen3.6-35B-A3B-Thinking` | 38.8 | 24k | MTP, vision, reasoning | hard problems, slow-but-smart answers |
| `gpt-oss-20b` | 38.4 | **64k** | reasoning, tools | fast reasoning + tool use · fast long-context (37 t/s vs Coder-Next 16; 131k ruled out) |
| `gpt-oss-20b-low` | 37.2 | **64k** | reasoning LOW, TTFT 0.8 s | same model, snappy answers (no long preamble) |
| `Qwen3.6-35B-A3B-MTP-UD-IQ3_XXS` | 35.9 | 24k | MTP, vision, thinking OFF | **daily driver** — compact instant answers |
| `gemma-4-26B-A4B-it-UD-IQ4_XS` | 33.9 | 24k | vision, MTP | quality generalist + best vision |
| `Qwen3.5-9B-UD-Q6_K_XL` | 32.1 | 32k | vision, MTP, reasoning | small Qwen, quick tasks |
| `Qwen3-Coder-30B-Instruct-UD-Q3_K_XL` | **~30** clean / 27.0 evening @moe12 | 32k | — | **main agent coder** — won the 2026-08-06 quant race; moe12 verified evening 08-07 + clean 08-08 |
| `GLM-4.7-Flash-UD-Q4_K_XL` | **21.5** @moe22 | 24k | reasoning | quality coder (opencode subagents); KV fix + moe22 ✓ verified clean 08-08 |
| `Qwen3-Coder-Next-UD-IQ3_XXS` | 16.0 | **128k** | 80B-A3B | long-session coder (128k ctx) |
| `Qwen3-Embedding-0.6B` | 33 (CPU) | 8k | CPU-only, `/v1/embeddings` | RAG/search embedder (not a chat model) |
Expected run-to-run spread: MTP models swing ±15% with draft **acceptance rate** (content-
dependent); CPU-heavy presets (Coder-Next, embedder) dip under daytime CPU contention.
Treat clean night runs as the reference; don't retune on daytime deltas.
## Which model, when
- **Agent coding loops** (edit/test cycles): `Qwen3-Coder-30B` — best speed/quality balance at 32k.
- **Hard code, reviews, tricky bugs**: `GLM-4.7-Flash` — strongest 30B-class coder, slightly slower.
- **Marathon sessions / huge conversation history**: `Qwen3-Coder-Next` — 128k ctx at only
6 GB VRAM (hybrid attention). Caveat: ~33 t/s prefill means it's for *growing* sessions
(`cache-reuse` makes turns incremental), **not** for cold-dumping 100k tokens.
- **Everyday questions**: `Qwen3.6-35B-A3B` — no reasoning preamble, 36 t/s.
- **Hard reasoning**: `35B-Thinking` (quality) or `gpt-oss-20b` (speed + tool use).
Their multi-second TTFT is the reasoning phase streaming first — not a slow load.
For quick interactive gpt-oss answers use `gpt-oss-20b-low` (reasoning effort low).
- **Images**: `gemma-4-26B` for quality, `gemma-4-E4B` or `Qwen3.5-9B` for speed,
`Qwen3.6-35B` when you want the daily driver to see the screenshot.
- **Long one-shot documents**: `gemma-4-E4B` — 185 t/s prefill eats 64k in ~6 min
(Coder-Next would take ~30 min to prefill the same).
- **RAG embeddings**: point Open WebUI etc. at `Qwen3-Embedding-0.6B`. Runs on CPU by
design — 0 VRAM, can never contribute to an overcommit.
## Speech-to-text (separate service)
STT deliberately does **not** live in this router — it is a CPU-only whisper.cpp service on port
11345, documented in `../whisper/README.md`. Two reasons: `models-max 1` would make every
transcription LRU-evict the resident chat model, and llama.cpp decodes audio via miniaudio
(wav/mp3/flac only) while browsers record opus.
The router *does* answer `POST /v1/audio/transcriptions` — it rewrites the request into a chat
completion — so `gemma-4-E4B-it-UD-Q8_K_XL` (audio-capable via its mmproj) transcribes uploaded
wav/mp3/flac with no config change. Keep that as the fallback; mind the eviction.
## Text-to-speech (separate service)
TTS is **not servable from this router at all** — unlike STT, there is no fallback. `llama-server`
has `--model-vocoder` / `--tts-use-guide-tokens`, but exposes **no `/v1/audio/speech` route**:
the router 404s on it exactly like an undefined path, and `grep -rn "audio/speech"` over the
b10216 tree returns nothing — the only audio route registered is `/v1/audio/transcriptions` (STT).
Worse, `--model-vocoder` is accepted by the server and then **silently ignored**: it is registered
in the arg table and documented, but no server code path consumes it. `--help` listing it is not
evidence it works. `llama-tts` is a CLI. (All verified 2026-08-09.)
So TTS is a CPU-only Kokoro-82M service on port **11347**, documented in `../tts/README.md`
(repo-side only — not installed on fl yet). The `models-max 1` argument applies even harder there
than for STT: TTS fires on *every* response, so a resident preset would evict the chat model every
single turn.
## Client wiring
The same catalog is mirrored in every client — when adding/removing a preset, update all:
- `common/.config/opencode/opencode.json` (both providers + agent model overrides)
- `common/.pi/agent/models.json` (both providers; `contextWindow` = server `ctx-size`)
- `common/.config/aichat/config.yaml` (both clients)
Rule: client model **id = config.ini section name**, client context ≤ server `ctx-size`.
The embedder is deliberately absent from chat clients.
2026-08-06: the coder id changed `…-IQ4_XS``…-UD-Q3_K_XL` in all six client files
(common + lw overlays) after the quant race; OWUI picks the new id up automatically from
the router, but chats/presets saved against the old id need re-picking.
## Tuning cheat-sheet
- **`n-cpu-moe`** (MoE only): experts→CPU. The speed/VRAM dial. Measured curve is gentle —
tune in steps of 24 layers, re-bench, keep free ≥2.5 GB.
- **`ctx-size`**: KV cost varies wildly by arch — Coder-Next 64k→128k cost 0.5 GB (hybrid
attn); dense Qwens pay ~1 GB per 16k. Raise only after a bench shows the headroom.
- **Quants**: dense = bandwidth-bound, quant size sets speed directly (9B: Q8→Q6 = +87%
with MTP). MoE = buy quality with a bigger quant, pay in CPU offload (both coders run
4-bit now; 3-bit was only ~2 t/s faster). Coders want ≥4-bit; chat tolerates UD 3-bit.
- **MTP / spec decode** (`spec-type draft-mtp`): ~1.52× decode. Qwen3.5/3.6 embed the
head in the *MTP-repo* GGUF (same filename as plain repo — size is the tell);
gemma-4 uses a separate `mtp-*.gguf` draft file. **No MTP possible yet for**:
GLM-4.7-Flash (conversion drops the head — expected ~1.5× when llama.cpp lands it),
Qwen3-Coder-Next (Qwen3-Next head exists upstream, no GGUF ships it), gpt-oss and
Qwen3-Coder-30B (no head exists). Recheck releases occasionally.
- **KV cache**: `q8_0` K / `q4_0` V everywhere except gpt-oss (attention-sink issues with
quantized KV — keep defaults there).
## Benchmarking
```
./bench.py # full sweep (add -x Qwen3-Embedding-0.6B — it can't chat)
./bench.py -m id1,id2 # just the changed presets
./bench.py -n 512 --ctx 16000 # longer gen + long-context decode column
```
**Benches need the server to themselves**: with `models-max 1`, a request for another
model arriving mid-load LRU **force-kills the loading instance** → fake "failed to load"
rows (and a dirty baseline). Close chat clients / OWUI before sweeping.
Safety-first: force-unloads resident models via the router API (`POST /models/unload`),
verifies a clean card with rocm-smi before every load, records a ⚠ row and skips
generation when a model lands too tight (free < 1.5 GB or GTT ballooning).
Results: `bench-results.md` (latest) + `bench-history.md` (append-only; diff runs there).
## Watch list
- **GLM MTP in llama.cpp** — the single biggest pending win (primary coder ~22 → ~30+).
- **Coder-Next 256k**: if the load log's `n_ctx_train` says 262144, ctx can likely double
again for ~1 GB (KV measured nearly flat).