Files
dots/lw/.config/llamacpp/README.md
T
2026-08-13 01:11:41 +02:00

134 lines
8.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Local LLM server — laptop (lw), setup & usage
llama.cpp in **router mode** on the laptop. This is the **CPU** counterpart of the fl box
(`fl/.config/llamacpp/`) — lw has **no discrete GPU**, so everything runs on the CPU and
the only resource ceiling is **system RAM** (not VRAM). None of fl's VRAM/freeze rules apply.
- **Box**: Intel Core i7-5600U (Broadwell 2015, **2 cores / 4 threads**, AVX2) · iGPU HD Graphics
5500 (no compute) · **15.49 GiB RAM** · DDR3L-1600 dual-channel: ~20 GB/s theoretical, **~810
GB/s effective measured** — the dense-decode bottleneck → prefer **small quants + MoE**.
- **Endpoint**: `http://127.0.0.1:11343/v1`**local only** (server binds `127.0.0.1`; deliberate:
the laptop roams onto untrusted wifi and the API has no key, unlike the stationary fl box).
- **Files here** (`~/.config/llamacpp/`, stowed from the repo's `lw/.config/llamacpp/`):
`config.ini` (presets — section names ARE the API model ids), `bench.py` (CPU benchmark),
`bench-results.md` / `bench-history.md` (ledgers, generated by bench.py). The GGUF blobs
themselves live in `~/software/models/` and are not tracked.
## How it runs
Launched by the fish abbr **`llamaserver`** (defined in `lw/.config/fish/host.fish`;
llama.cpp is pacman-installed, `llama-server` is on PATH):
```
nice -n 19 llama-server --host 127.0.0.1 --port 11343 --cors-origins localhost \
--models-max 1 --models-preset ~/.config/llamacpp/config.ini
```
- Router mode: each requested model loads on demand in a child process.
- **`--models-max 1`**: one model resident; a new request LRU-evicts the previous. There's no
fl-style freeze risk here (no display-driving GPU) — this is purely the safest RAM posture on
15 GB with no co-residency swap. Raise to 23 only to keep small models hot together
(E2B+4B+1.7B ≈ 6 GB fits). `sleep-idle-seconds` per preset frees idle models either way.
- **`nice -n 19`**: on 2 cores generation pins the whole CPU — niceness makes inference yield to
interactive apps (full speed when the desktop is idle).
- **`--cors-origins localhost`**: same-machine profile from llama.cpp PR 25655 — the default CORS
reflects any Origin, so without this a malicious page in the laptop's own browser could call
the (keyless) API on 127.0.0.1. Browser-only enforcement: bench.py/CLI clients are unaffected.
- `config.ini` here is the stowed repo copy (`lw/.config/llamacpp/`), so editing it edits the
repo (no deploy step).
## RAM (the golden rule)
There is no VRAM and no freeze risk. The constraint is plain: **N resident models must fit in
system RAM together.** `bench.py` records each model's RAM footprint (MemAvailable delta) and
the remaining free RAM — keep enough headroom for the desktop (and for co-resident models if you
ever raise `--models-max` above 1). If RAM gets tight: shorten `sleep-idle-seconds`, use smaller
quants, or reduce `ctx-size`. Note: the 4B/E2B presets are small enough to co-reside, but a big
MoE (e.g. Qwen3-30B-A3B at ~12 GB) fills most of 15 GB → it runs **1-at-a-time** regardless.
`threads = 2` (physical cores) on the dense presets — they measure at the practical bandwidth
wall (~810 GB/s effective) and the 2026-07-26 sweep confirmed more threads don't lift them.
The **compute-bound** pair runs `threads = 3` (same-run sweep: Granite 4.5→4.8, 30B 1.9→2.0;
`threads = 4` LOSES to HT contention on both). Every preset sets `threads-batch = 4`: prefill
is compute-bound on all models — benched +8% prefill / 9% TTFT with no decode cost, and the
extra HT load only bursts during prefill.
## Tuning cheat-sheet (CPU)
- **Quant size is the speed dial**: decode is bandwidth-bound, so a smaller quant is directly
faster. Prefer Q4_K_M / Q5_K_M; F16/Q8 are usually quality-overkill AND slow on CPU.
- **`threads`** = **physical** core count (`lscpu` → Core(s)/socket × Socket(s)). Hyperthreads
rarely help decode; oversubscribing can hurt. Set this per preset (default 4 is a guess).
- **Model size**: stick to ≤ ~8B. **MoE** models (few active params, e.g. Qwen3-A3B) are ideal
IF they fit RAM — they decode at their *active* size while giving bigger-model quality. There
is no `n-cpu-moe` tradeoff here — it's all CPU already.
- **MTP / speculative decode** (`spec-type = draft-mtp`): a clear win on the fl GPU (+50100%) but
**TESTED AND REJECTED on this box** (2026-07-26): the draft lost the final CPU-pinned
(`gpu-layers-draft = 0`) same-run A/Bs on both gemmas — E2B 4.0 vs 4.9 t/s at 47% acceptance,
E4B 2.3 vs 2.4 at 62%. Batched verify triples the FLOPs per weight-read and 2 cores have
no spare compute, so spec decode can't pay here at any setting. MTP stays an fl-only trick.
(Only gemma-4 has drafts anyway; Qwen 2507/2.5 have no MTP path — that needs Qwen 3.5/3.6.)
- **KV cache**: leave at default (f16). Quantized KV (`cache-type-*`) adds dequant overhead with
no VRAM to reclaim → usually slower on CPU. (This is the opposite of the fl config.)
- **iGPU / Vulkan**: the pacman build **does** have a Vulkan backend and by default auto-offloads
layers to the HD 5500 — that's why every preset pins `n-gpu-layers = 0`. The iGPU shares the
same DDR3, is slower than the cores, and big allocations die (`ErrorOutOfDeviceMemory`, SIGSEGV
in llama-cli) — it broke the 30B load outright. Only re-enable offload if a bench proves the
iGPU wins (it won't).
## Benchmarking
Run **on the laptop** (against its own localhost server):
```
./bench.py # full sweep of everything the server lists
./bench.py -m id1,id2 # just the changed presets
./bench.py -n 512 --ctx 8000 # longer gen + a long-context decode column
```
For a clean per-model RAM footprint, restart the server first (with `--models-max 1` only a
re-bench of the *same* still-resident model reads a `~0` RAM Δ, but a restart keeps runs comparable). Results land in
`bench-results.md` (latest) + `bench-history.md` (append-only — diff runs there).
⚠ Absolute numbers swing **±2030%** with desktop load and thermals (E2B has measured 6.1 idle
vs 4.9 warm/busy). Only rows from the **same run** are directly comparable — A/B via extra
preset ids in one `./bench.py -m a,b` command (see the rig sections in config.ini), never by
comparing across runs. For spec/MTP rows the draft **acceptance %** appears in the spec column.
## Model roster (reference numbers = clean sweep 2026-07-25, n-gpu-layers=0; Granite/30B since
promoted to threads=3 and all presets to threads-batch=4, 2026-07-26)
| model id | decode t/s | prefill t/s | ctx | extras | role |
|---|---|---|---|---|---|
| `gemma-4-E2B-it-UD-Q4_K_XL` | **6.1** | 8 | 4k | vision-capable (mmproj) | fastest — quick Q&A, askllama-class |
| `Qwen3-1.7B` | 4.8 | 11 | 8k | thinking | snappy small tasks |
| `Granite-4.0-H-Tiny` | 4.34.8 | 1114 | 8k | 7B/~1B-active hybrid MoE, threads 3 | ⭐ speed AND brains |
| `Qwen3-4B-Instruct-2507` | 3.9 | 7 | 4k | non-thinking | daily driver |
| `Qwen2.5-Coder-3B-Instruct` | 3.3 | 6 | 8k | temp 0.2 | small coder |
| `Qwen3-30B-A3B-Instruct-2507-UD-IQ3_XXS` | 1.92.7 | 4 | 4k | ~12 GB, TTFT ~6 s, threads 3 | ⭐ quality when you can wait |
| `gemma-4-E4B-it-UD-Q4_K_XL` | 2.42.5 | 4 | 4k | vision (mmproj) | quality vision; else the 30B is smarter at the same speed |
| `Jan-v3-4b` | 2.1¹ | 4 | 4k | agentic tune | agentic variant of the 4B (~0.9× its speed) |
¹ variance-depressed in the sweep; same-run A/B vs the base (2026-07-26) shows ~0.9×.
## Client wiring
Clients run **on the laptop itself** (the server is localhost-only, see Endpoint above); same
rule as fl otherwise: client model **id = config.ini section name**, client context ≤ server
`ctx-size`. Point it at `http://127.0.0.1:11343/v1` (no key). If LAN access is ever needed,
switch the abbr back to `--host 0.0.0.0` **and add `--api-key`** — never open it keyless.
**Wired (2026-07-26), lw-only via stow host overlays** — these are full-file `--override`
copies of the common configs with an added local provider; when the common client configs
change, mirror the change here:
| tool | overlay file | provider / client id |
|---|---|---|
| aichat | `lw/.config/aichat/config.yaml` | `lw:` (e.g. `lw:gemma-4-E2B-it-UD-Q4_K_XL`) |
| opencode | `lw/.config/opencode/opencode.json` | `lwcpp/` |
| pi | `lw/.pi/agent/models.json` + `settings.json` | `lwcpp` (⚠ pi hides providers not in settings.json `enabledModels`) |
All 8 presets are listed in each, with client context capped to the server `ctx-size` and the
benched t/s in the display names. Defaults still point at the fl LAN models — switch to a local
model in-tool when roaming. Re-run `./install.sh` on lw once so stow links the new files.