Files
dots/lw/.config/llamacpp/config.ini
T
coja 1e39e0cb4b [Sync] replace the flat layout with the unified stow tree
Supersedes the old flat .config/ layout (last published 2026-06-28) with the
private repo's structure: one shared base plus per-host overlays.

- packages: common/ gui/ lw/ fl/ wm/ plus install.sh and bin/ tooling
  (dotsync, reconcile-hyde.sh)
- new README covering the layout, deploy order and the HyDE dependency
- current HyDE waybar rig (layouts/, cava), pi agent extensions, claude/
  config, tmux, presenterm, aichat roles
- fish: kp (keepassxc-cli + fzf picker, db path from $KP_DB) and
  bind_M_n_history (alt+1..9 recalls the nth history entry)
- drops cruft that should never have been tracked: the duplicate top-level
  .pi/ copy, btop.log, zellij config.kdl.bak, fish_variables
- .pi/agent/auth.json is gitignored; auth.json.example ships instead

Host-specific work sessions and the personal backlog stay in the private
tree. Endpoint locators in the llamacpp/whisper guides are placeholders
($SERVER, <own-domain>) — the guides themselves stay, since they are the
useful part.
2026-08-13 02:53:06 +02:00

215 lines
13 KiB
INI
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# llama.cpp model config — LAPTOP (lw) · CPU inference, NO discrete GPU
# ═════════════════════════════════════════════════════════════════════════════
# HARDWARE: Intel Core i7-5600U (Broadwell 2015, 2 cores / 4 threads, AVX2, no AVX-512)
# · iGPU HD Graphics 5500 (i915, not used for compute) · 15.49 GiB RAM
# · DDR3L-1600 dual-channel ≈ ~20 GB/s effective → THE bottleneck.
#
# This is the laptop analog of fl/.config/llamacpp/config.ini (the RX 7600 XT box), but on a
# CPU-only machine almost every fl rule inverts:
#
# • NO VRAM / GTT / freeze risk — no display-driving GPU to starve. The only ceiling is SYSTEM
# RAM. The `llamaserver` abbr uses `--models-max 1` (one model resident; a new request LRU-
# evicts the previous — safest on this 15 GB laptop, no co-residency swap). Raise to 2-3 only
# to keep small models hot together (E2B+4B+1.7B ≈ 6 GB fits).
# • Decode speed ≈ effective bandwidth ÷ active-bytes, so SMALL QUANTS are directly faster and
# MoE (few active params) is the single biggest trick. Clean bench 2026-07-25 (n-gpu-layers=0):
# dense models reach ~8-10 GB/s effective (decode t/s × model GB) ≈ the PRACTICAL DDR3L wall,
# i.e. dense decode IS bandwidth-bound as designed. (The earlier "core-starved ~5-7 GB/s"
# reading was the Vulkan-offload taint.) Granite (~2.6 GB/s eff) and the 30B (~3.5) sit far
# BELOW the wall → COMPUTE-bound: threads 3/4 might lift those two (dense won't gain).
# • `threads = 2` = PHYSICAL cores for the dense presets (at the wall — swept 2026-07-26, no
# gain). `threads = 3` on the compute-bound Granite/30B: +~5-7% same-run (t4 LOSES to HT
# contention on both — never promote it). `threads-batch = 4` EVERYWHERE: prefill is
# compute-bound on all models — benched +8% prefill / 9% TTFT on Granite, no decode cost
# (HT only bursts during prefill, so no sustained desktop pressure).
# • ⚠ RESPONSIVENESS: on 2 cores, generation pins the whole CPU → the desktop can freeze. The
# `llamaserver` abbr runs the server under `nice -n 19`, so inference YIELDS to interactive apps
# (still full speed when the desktop is idle). Still laggy? set `threads = 1` (≈half speed, always
# responsive).
# • `flash-attn = on` EVERYWHERE EXCEPT the two gemma-4 presets: benched ~+11% here (gen 3.7→4.1,
# prompt 1.7→1.9 t/s) so it's worth keeping — BUT this build (b10068) FAILS TO LOAD the gemma3n
# arch (E2B/E4B) with flash-attn, so those two stay OFF. (If another model also won't load with
# it, drop it there too.) No cache-type-k/v (quantized KV) on CPU: dequant overhead, no VRAM to
# reclaim → slower. KV stays f16.
# • ⚠ `n-gpu-layers = 0` REQUIRED in every preset — the pacman build has a VULKAN backend and
# by default AUTO-OFFLOADS layers to the HD 5500 iGPU (common_fit_params). Seen live
# 2026-07-25: Granite/30B loads died with "ggml_vulkan ... ErrorOutOfDeviceMemory" (llama-cli
# even SIGSEGVs). The iGPU shares the same DDR3 and is slower than the cores — offload also
# silently poisons benchmarks ("omit it for CPU-only" was wrong on this build).
#
# Clean baseline 2026-07-25 17:41 (n-gpu-layers=0, threads=2, gen 256 — bench-history.md);
# decode noted per preset below, roster table in the README. Prefill is slow on 2 cores, so
# huge ctx is slow to FILL (not RAM-limited) — 4096 is a sane default; raise only if needed.
#
# Router: llama-server --host 127.0.0.1 --port 11343 --cors-origins localhost --models-max 1 \
# (localhost-only: the laptop roams onto untrusted networks and there's no API key —
# unlike fl, nothing on the LAN targets this server. --cors-origins localhost blocks
# browser pages from calling the keyless API; non-browser clients unaffected — PR 25655)
# --models-preset ~/.config/llamacpp/config.ini (abbr `llamaserver`; llama.cpp is
# pacman-installed → `llama-server` is on PATH)
# This file stows to ~/.config/llamacpp/ (like fl); the GGUFs live in ~/software/models/.
# Section names ARE the API model ids.
# ═════════════════════════════════════════════════════════════════════════════
# ─────────────────────────────────────────────────────────────────────────────
# Fast / small — the "instant" tier (quick Q&A, askllama, autocomplete)
# ─────────────────────────────────────────────────────────────────────────────
[gemma-4-E2B-it-UD-Q4_K_XL]
# Gemma 3n E2B (~2B effective/active) — FASTEST here: benched 6.1 t/s (clean 2026-07-25). Add an mmproj for vision.
model = /home/anon/software/models/gemma-4-E2B-it-UD-Q4_K_XL.gguf
# flash-attn OFF: gemma3n fails to load with it on this build (b10068).
ctx-size = 4096
threads = 2
threads-batch = 4
n-gpu-layers = 0
jinja = on
temp = 1.0 # Gemma default (1.0 / top-p 0.95 / top-k 64); drop to 0.7 for determinism
top-p = 0.95
top-k = 64
sleep-idle-seconds = 30
[Qwen3-1.7B]
# Qwen3-1.7B, Q8_0 (~1.8 GB — Q8 is fine at this size). Benched 4.8 t/s (clean 2026-07-25) —
# ~8.8 GB/s effective ≈ at the bandwidth wall; more threads won't lift this one.
model = /home/anon/software/models/Qwen3-1.7B-Q8_0.gguf
ctx-size = 8192 # tiny KV — can afford more ctx
threads = 2
threads-batch = 4
n-gpu-layers = 0
flash-attn = on
jinja = on
temp = 0.7
top-p = 0.8
top-k = 20
min-p = 0
repeat-penalty = 1.05
sleep-idle-seconds = 30
# ─────────────────────────────────────────────────────────────────────────────
# Daily drivers — 4B-class (best quality/speed balance; benched 2.1-3.9 t/s, clean 2026-07-25)
# ─────────────────────────────────────────────────────────────────────────────
[Qwen3-4B-Instruct-2507]
# Qwen3-4B Instruct 2507, Q4_K_M (~2.5 GB). Benched 3.9 t/s (clean 2026-07-25). Best small
# all-rounder; non-thinking, compact answers. The natural DAILY DRIVER — askllama points here.
model = /home/anon/software/models/Qwen3-4B-Instruct-2507.Q4_K_M.gguf
ctx-size = 4096
threads = 2
threads-batch = 4
n-gpu-layers = 0
flash-attn = on
jinja = on
temp = 0.7
top-p = 0.8
top-k = 20
min-p = 0
repeat-penalty = 1.05
sleep-idle-seconds = 30
[gemma-4-E4B-it-UD-Q4_K_XL]
# Gemma 3n E4B (~4B active), QAT Q4_K_XL (~4-7 GB). Quality step up from E2B; add an mmproj for vision.
# (File is the -qat- build — QAT keeps 4-bit quality; the id drops "qat" to match fl's convention.)
model = /home/anon/software/models/gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf
# flash-attn OFF: gemma3n fails to load with it on this build (b10068).
# Benched 2.4 t/s (nodraft, clean 2026-07-26) — same speed as the far-smarter 30B MoE, so E4B's
# niche is vision (mmproj). MTP draft TESTED AND REMOVED 2026-07-26: lost the final CPU-pinned
# same-run A/Bs on both gemmas (E2B 4.0 vs 4.9 at 47% acc; E4B 2.3 vs 2.4 at 62% acc) — batched
# verify is compute-bound on 2 cores; fl's +50% is a GPU result. The mtp-*.gguf drafts in
# ~/software/models/ can be deleted.
ctx-size = 4096
threads = 2
threads-batch = 4
n-gpu-layers = 0
jinja = on
temp = 1.0
top-p = 0.95
top-k = 64
sleep-idle-seconds = 30
[Jan-v3-4b]
# Jan v3 4B (Qwen3-4B-based, tuned for agentic/reasoning use). Same-run A/B vs its base
# (2026-07-26): 2.6 vs 2.9 t/s — NOT anomalous, just ~10% slower (Q4_K_XL reads a bit more per
# token than the base's Q4_K_M). The earlier 2.1-vs-3.9 reading was cross-run variance.
model = /home/anon/software/models/Jan-v3-4b-base-instruct-Q4_K_XL.gguf
ctx-size = 4096 # init had 256 — unusably small for reasoning; raised
threads = 2
threads-batch = 4
n-gpu-layers = 0
flash-attn = on
jinja = on
temp = 0.7
top-p = 0.95
top-k = 20
sleep-idle-seconds = 30
# ─────────────────────────────────────────────────────────────────────────────
# Coding
# ─────────────────────────────────────────────────────────────────────────────
[Qwen2.5-Coder-3B-Instruct]
# Qwen2.5-Coder-3B Instruct, Q5_K_M (~2.3 GB, official Qwen GGUF). Small dedicated coder — benched 3.3 t/s (clean 2026-07-25).
model = /home/anon/software/models/qwen2.5-coder-3b-instruct-q5_k_m.gguf
ctx-size = 8192
threads = 2
threads-batch = 4
n-gpu-layers = 0
flash-attn = on
jinja = on
temp = 0.2 # low temp for code
top-p = 0.9
sleep-idle-seconds = 30
# ─────────────────────────────────────────────────────────────────────────────
# ⭐ Quality when you can wait — 30B MoE, only 3B ACTIVE → ~30B quality at ~4B speed
# ─────────────────────────────────────────────────────────────────────────────
[Qwen3-30B-A3B-Instruct-2507-UD-IQ3_XXS]
# Qwen3-30B-A3B Instruct 2507, UD-IQ3_XXS (~12 GB). Benched 1.9-2.7 t/s decode (cross-run
# variance) / 4 t/s prefill / TTFT ~6.5 s — ~30B quality at E4B speed: THE quality pick when you
# can wait. Compute-bound (~3.5 GB/s eff) → thread sweep 2026-07-26: t3 2.0 > t2 1.9 > t4 1.8
# same-run → threads=3 promoted. (Earlier load failures were the Vulkan auto-offload OOM.)
# ⚠ RAM: ~12 GB on 15.5 GB. With `--models-max 1` (the abbr default) it's always alone — no
# co-residency swap risk — but 12 GB + desktop is still tight, so keep ctx small and close the
# browser. If it swaps even alone, step the quant down to UD-IQ2_M (~10.5 GB).
model = /home/anon/software/models/Qwen3-30B-A3B-Instruct-2507-UD-IQ3_XXS.gguf
ctx-size = 4096
threads = 3
threads-batch = 4
n-gpu-layers = 0
flash-attn = on
jinja = on
temp = 0.7
top-p = 0.8
top-k = 20
min-p = 0
sleep-idle-seconds = 30
# ─────────────────────────────────────────────────────────────────────────────
# ⭐ Faster "speed AND brains" — small-active MoE
# ─────────────────────────────────────────────────────────────────────────────
[Granite-4.0-H-Tiny]
# IBM Granite 4.0 H Tiny — 7B total / ~1B ACTIVE (hybrid Mamba-2 + MoE). Benched 4.3-4.8 t/s
# decode / 11-14 t/s prefill: faster than every 4B AND ~7B-class quality, tiny RAM Δ (mmap).
# COMPUTE-bound (~2.6 GB/s eff) → thread sweep 2026-07-26: t3 4.8 > t2 4.5 > t4 4.3 same-run →
# threads=3 promoted (t4 loses to HT contention). Revert to 2 if the desktop drags.
# QUANT: Q4_K_M (~4.2 GB) — goal here is speed, and Q4_K_M is the sweet spot (Q5_K_M/Q6_K fine
# too, RAM allows, small speed cost since decode reads only the ~1B active).
# Needs a current llama.cpp (hybrid arch) — the pacman build should be fine.
model = /home/anon/software/models/granite-4.0-h-tiny-Q4_K_M.gguf
ctx-size = 8192
threads = 3
threads-batch = 4
n-gpu-layers = 0
flash-attn = on # confirmed loading fine on the hybrid arch (2026-07-25)
jinja = on
temp = 0.7
top-p = 0.95
top-k = 20
sleep-idle-seconds = 30
# (Bench-rig history: the MTP A/B rigs and the thread-sweep rigs that lived here were deleted
# 2026-07-26 after their verdicts — MTP rejected (E4B comment), threads=3 promoted to
# Granite/30B, threads-batch=4 everywhere. Raw legs in bench-history.md.)