Files
dots/fl/.config/llamacpp/README.md
T
coja 1e39e0cb4b [Sync] replace the flat layout with the unified stow tree
Supersedes the old flat .config/ layout (last published 2026-06-28) with the
private repo's structure: one shared base plus per-host overlays.

- packages: common/ gui/ lw/ fl/ wm/ plus install.sh and bin/ tooling
  (dotsync, reconcile-hyde.sh)
- new README covering the layout, deploy order and the HyDE dependency
- current HyDE waybar rig (layouts/, cava), pi agent extensions, claude/
  config, tmux, presenterm, aichat roles
- fish: kp (keepassxc-cli + fzf picker, db path from $KP_DB) and
  bind_M_n_history (alt+1..9 recalls the nth history entry)
- drops cruft that should never have been tracked: the duplicate top-level
  .pi/ copy, btop.log, zellij config.kdl.bak, fish_variables
- .pi/agent/auth.json is gitignored; auth.json.example ships instead

Host-specific work sessions and the personal backlog stay in the private
tree. Endpoint locators in the llamacpp/whisper guides are placeholders
($SERVER, <own-domain>) — the guides themselves stay, since they are the
useful part.
2026-08-13 02:53:06 +02:00

167 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Local LLM server — setup & usage
llama.cpp in **router mode** serving ~10 model presets from one 16 GB GPU.
Last full tune: **2026-07-16** (decode numbers measured then via `bench.py`, unless a preset comment in `config.ini` says otherwise).
Last full clean sweep: **2026-08-08 03:22** — all 11 presets loaded and ran green (baseline 2.0 GB) after the week's coder-race/KV-fix/gemma-sidecar round.
- **Box**: AMD RX 7600 XT 16 GB (ROCm, ~288 GB/s) + Ryzen 5600X (6 cores) + 48 GB RAM.
⚠ The GPU **also drives the display** — see [VRAM safety](#vram-safety-the-golden-rules).
- **Endpoints**: LAN `http://$SERVER:11343/v1` · remote `https://<own-domain>/api/v1`
(own reverse proxy to the same server; requires an API key. Only the LAN endpoint is keyless).
`$SERVER` = the box's LAN address.
- **Files here**: `config.ini` (model presets — the section names ARE the API model ids),
`bench.py` (safety-first benchmark), `bench-results.md` / `bench-history.md` (ledgers, generated).
## How it runs
systemd unit **`llama.service`** (on the server, under a dedicated service user) runs a script wrapping:
```
llama-server --port 11343 --host 0.0.0.0 --models-max 1 --models-preset ~/.config/llamacpp/config.ini
```
- Router mode: each requested model loads on demand in a child process.
- **`--models-max 1` is deliberate and load-bearing**: at most one model resident; requesting
another LRU-evicts the current one *first*. Anything ≥2 lets two ~13 GB models stack →
VRAM overcommit → GTT spill → **whole-PC freeze**. Do not raise it casually.
(`sleep-idle-seconds` per preset is the second line of defense.)
- Models live in `~/software/models/` (paths in `config.ini` must stay absolute — llama-server
does not expand `~` or env vars in preset values). mmproj/draft files must match the exact
filename in the preset (watch for trailing spaces when renaming downloads!).
**Deploying config changes**: this dir is stowed into `~/.config/llamacpp/` as per-file
symlinks, so on the server `dotsync` (pull; re-stows only if files were added/removed) then
`sudo systemctl restart llama.service` — the router reads presets once at startup.
After ANY preset change: restart + `./bench.py -m <changed-ids>` and check the guards.
## VRAM safety (the golden rules)
ROCm does not OOM cleanly — an overcommitted model spills into GTT (system RAM),
starves the desktop and freezes the PC (reboot). Hard-won rules:
1. Keep every preset's benched **VRAM free ≥ 2.5 GB** (idle desktop uses ~1.3 of 17.2).
2. File size > VRAM is **fine** for MoE — `n-cpu-moe N` keeps the first N layers' experts
in system RAM. It is the main tuning knob: higher = safer/slower, lower = faster.
3. Bench with a **clean baseline** (close the browser): a dirty baseline both skews the
margins and costs ~34 t/s of decode (GPU compute contention — measured 2026-07-08).
4. Bench after every change; config-edit without restart = numbers from the *old* config.
## Model roster (reference numbers = clean night sweeps, 2026-07-16)
| model id | decode t/s | ctx | extras | role |
|---|---|---|---|---|
| `gemma-4-E4B-it-UD-Q8_K_XL` | **57.3** | **96k** | vision, MTP, prefill 185 t/s | fast small generalist, long docs, images (ctx 96k ✓ 08-08, n_ctx_train 131072) |
| `Qwen3.6-35B-A3B-Thinking` | 38.8 | 24k | MTP, vision, reasoning | hard problems, slow-but-smart answers |
| `gpt-oss-20b` | 38.4 | **64k** | reasoning, tools | fast reasoning + tool use · fast long-context (37 t/s vs Coder-Next 16; 131k ruled out) |
| `gpt-oss-20b-low` | 37.2 | **64k** | reasoning LOW, TTFT 0.8 s | same model, snappy answers (no long preamble) |
| `Qwen3.6-35B-A3B-MTP-UD-IQ3_XXS` | 35.9 | 24k | MTP, vision, thinking OFF | **daily driver** — compact instant answers |
| `gemma-4-26B-A4B-it-UD-IQ4_XS` | 33.9 | 24k | vision, MTP | quality generalist + best vision |
| `Qwen3.5-9B-UD-Q6_K_XL` | 32.1 | 32k | vision, MTP, reasoning | small Qwen, quick tasks |
| `Qwen3-Coder-30B-Instruct-UD-Q3_K_XL` | **~30** clean / 27.0 evening @moe12 | 32k | — | **main agent coder** — won the 2026-08-06 quant race; moe12 verified evening 08-07 + clean 08-08 |
| `GLM-4.7-Flash-UD-Q4_K_XL` | **21.5** @moe22 | 24k | reasoning | quality coder (opencode subagents); KV fix + moe22 ✓ verified clean 08-08 |
| `Qwen3-Coder-Next-UD-IQ3_XXS` | 16.0 | **128k** | 80B-A3B | long-session coder (128k ctx) |
| `Qwen3-Embedding-0.6B` | 33 (CPU) | 8k | CPU-only, `/v1/embeddings` | RAG/search embedder (not a chat model) |
Expected run-to-run spread: MTP models swing ±15% with draft **acceptance rate** (content-
dependent); CPU-heavy presets (Coder-Next, embedder) dip under daytime CPU contention.
Treat clean night runs as the reference; don't retune on daytime deltas.
## Which model, when
- **Agent coding loops** (edit/test cycles): `Qwen3-Coder-30B` — best speed/quality balance at 32k.
- **Hard code, reviews, tricky bugs**: `GLM-4.7-Flash` — strongest 30B-class coder, slightly slower.
- **Marathon sessions / huge conversation history**: `Qwen3-Coder-Next` — 128k ctx at only
6 GB VRAM (hybrid attention). Caveat: ~33 t/s prefill means it's for *growing* sessions
(`cache-reuse` makes turns incremental), **not** for cold-dumping 100k tokens.
- **Everyday questions**: `Qwen3.6-35B-A3B` — no reasoning preamble, 36 t/s.
- **Hard reasoning**: `35B-Thinking` (quality) or `gpt-oss-20b` (speed + tool use).
Their multi-second TTFT is the reasoning phase streaming first — not a slow load.
For quick interactive gpt-oss answers use `gpt-oss-20b-low` (reasoning effort low).
- **Images**: `gemma-4-26B` for quality, `gemma-4-E4B` or `Qwen3.5-9B` for speed,
`Qwen3.6-35B` when you want the daily driver to see the screenshot.
- **Long one-shot documents**: `gemma-4-E4B` — 185 t/s prefill eats 64k in ~6 min
(Coder-Next would take ~30 min to prefill the same).
- **RAG embeddings**: point Open WebUI etc. at `Qwen3-Embedding-0.6B`. Runs on CPU by
design — 0 VRAM, can never contribute to an overcommit.
## Speech-to-text (separate service)
STT deliberately does **not** live in this router — it is a CPU-only whisper.cpp service on port
11345, documented in `../whisper/README.md`. Two reasons: `models-max 1` would make every
transcription LRU-evict the resident chat model, and llama.cpp decodes audio via miniaudio
(wav/mp3/flac only) while browsers record opus.
The router *does* answer `POST /v1/audio/transcriptions` — it rewrites the request into a chat
completion — so `gemma-4-E4B-it-UD-Q8_K_XL` (audio-capable via its mmproj) transcribes uploaded
wav/mp3/flac with no config change. Keep that as the fallback; mind the eviction.
## Text-to-speech (separate service)
TTS is **not servable from this router at all** — unlike STT, there is no fallback. `llama-server`
has `--model-vocoder` / `--tts-use-guide-tokens`, but exposes **no `/v1/audio/speech` route**:
the router 404s on it exactly like an undefined path, and `grep -rn "audio/speech"` over the
b10216 tree returns nothing — the only audio route registered is `/v1/audio/transcriptions` (STT).
Worse, `--model-vocoder` is accepted by the server and then **silently ignored**: it is registered
in the arg table and documented, but no server code path consumes it. `--help` listing it is not
evidence it works. `llama-tts` is a CLI. (All verified 2026-08-09.)
So TTS is a CPU-only Kokoro-82M service on port **11347**, documented in `../tts/README.md`
(repo-side only — not installed on fl yet). The `models-max 1` argument applies even harder there
than for STT: TTS fires on *every* response, so a resident preset would evict the chat model every
single turn.
## Client wiring
The same catalog is mirrored in every client — when adding/removing a preset, update all:
- `common/.config/opencode/opencode.json` (both providers + agent model overrides)
- `common/.pi/agent/models.json` (both providers; `contextWindow` = server `ctx-size`)
- `common/.config/aichat/config.yaml` (both clients)
Rule: client model **id = config.ini section name**, client context ≤ server `ctx-size`.
The embedder is deliberately absent from chat clients.
2026-08-06: the coder id changed `…-IQ4_XS``…-UD-Q3_K_XL` in all six client files
(common + lw overlays) after the quant race; OWUI picks the new id up automatically from
the router, but chats/presets saved against the old id need re-picking.
## Tuning cheat-sheet
- **`n-cpu-moe`** (MoE only): experts→CPU. The speed/VRAM dial. Measured curve is gentle —
tune in steps of 24 layers, re-bench, keep free ≥2.5 GB.
- **`ctx-size`**: KV cost varies wildly by arch — Coder-Next 64k→128k cost 0.5 GB (hybrid
attn); dense Qwens pay ~1 GB per 16k. Raise only after a bench shows the headroom.
- **Quants**: dense = bandwidth-bound, quant size sets speed directly (9B: Q8→Q6 = +87%
with MTP). MoE = buy quality with a bigger quant, pay in CPU offload (both coders run
4-bit now; 3-bit was only ~2 t/s faster). Coders want ≥4-bit; chat tolerates UD 3-bit.
- **MTP / spec decode** (`spec-type draft-mtp`): ~1.52× decode. Qwen3.5/3.6 embed the
head in the *MTP-repo* GGUF (same filename as plain repo — size is the tell);
gemma-4 uses a separate `mtp-*.gguf` draft file. **No MTP possible yet for**:
GLM-4.7-Flash (conversion drops the head — expected ~1.5× when llama.cpp lands it),
Qwen3-Coder-Next (Qwen3-Next head exists upstream, no GGUF ships it), gpt-oss and
Qwen3-Coder-30B (no head exists). Recheck releases occasionally.
- **KV cache**: `q8_0` K / `q4_0` V everywhere except gpt-oss (attention-sink issues with
quantized KV — keep defaults there).
## Benchmarking
```
./bench.py # full sweep (add -x Qwen3-Embedding-0.6B — it can't chat)
./bench.py -m id1,id2 # just the changed presets
./bench.py -n 512 --ctx 16000 # longer gen + long-context decode column
```
**Benches need the server to themselves**: with `models-max 1`, a request for another
model arriving mid-load LRU **force-kills the loading instance** → fake "failed to load"
rows (and a dirty baseline). Close chat clients / OWUI before sweeping.
Safety-first: force-unloads resident models via the router API (`POST /models/unload`),
verifies a clean card with rocm-smi before every load, records a ⚠ row and skips
generation when a model lands too tight (free < 1.5 GB or GTT ballooning).
Results: `bench-results.md` (latest) + `bench-history.md` (append-only; diff runs there).
## Watch list
- **GLM MTP in llama.cpp** — the single biggest pending win (primary coder ~22 → ~30+).
- **Coder-Next 256k**: if the load log's `n_ctx_train` says 262144, ctx can likely double
again for ~1 GB (KV measured nearly flat).