# LLM inference backend for Hermes (valhalla) — design **Date:** 2026-06-26 **Status:** Approved (pending spec review) **Stack:** new `stacks/llm/` ## Goal Stand up a local, OpenAI-compatible LLM inference endpoint on valhalla and point the existing **Hermes** agent backend at it, so Hermes runs against a self-hosted 14B model instead of an external provider. ## Key decision: llama.cpp, not vLLM The request was "set up vLLM," but valhalla's GPU is a **Tesla P100-PCIE-16GB**, die **GP100, compute capability 6.0**. That rules out vLLM for the desired model class: - The GP100 (cc 6.0) lacks the **DP4A INT8** instructions that vLLM's AWQ/GPTQ kernels require (those need Turing 7.5 / Ampere). So vLLM **cannot run quantized 13B+** here. - An unquantized 13B in fp16 is ~26 GB → does not fit 16 GB. - Current vLLM refuses cc < 7.0 outright; even a ~7B fp16 would need a pinned old vLLM + `--dtype float16` + `VLLM_ATTENTION_BACKEND=XFORMERS`, and still no 13B. Since the user wants a **13B+ class model** and only needs an **OpenAI-compatible** API (confirmed), the right engine is **llama.cpp's `llama-server`**: rock-solid Pascal (sm_60) support, GGUF quantization, and a native OpenAI-compatible `/v1` API that Hermes consumes unchanged. ## Hardware / host facts (verified live 2026-06-26) - GPU: 1× Tesla P100-PCIE-16GB, cc 6.0, driver 580.159.04, CUDA 13. Idle. (2nd staged P100 not installed — single-GPU design.) - Docker 29.5.2, NVIDIA Container Toolkit 1.19.1. - **CDI already configured:** `nvidia.com/gpu=0` is a valid device. No host runtime changes needed — the compose just references the CDI device. - `/storage1` (ZFS, virtiofs): 17 TB free. Weights live here; VM root is only 200 GB. - **Hermes** runs on the host (not Docker) from `/home/ginnoir/.hermes/`: - `hermes dashboard --host 172.20.0.1 --port 9119` (gateway/agent API+WS, pid 1116) - `hermes_cli.main gateway run` (pid 1119) - `hermes-webui/server.py` on `172.20.0.1:8787` (pid 1128) - Bound to `172.20.0.1` = the Docker bridge gateway, so host↔container reachability is trivial: a container published on `172.20.0.1:` is reachable by Hermes and by other containers. No Caddy hop in the inference path. ## Architecture ``` Hermes agent/gateway (host, 172.20.0.1) │ OpenAI base_url → http://172.20.0.1:8090/v1 (api_key = LLM_API_KEY) ▼ llama-server container (stacks/llm) ──CDI nvidia.com/gpu=0──▶ Tesla P100 model: Qwen2.5-14B-Instruct-Q4_K_M.gguf weights bind-mounted from /storage1/labdata/llm/models ``` No Caddy endpoint (Hermes-only, per decision). Endpoint is unauthenticated-but-API-keyed and only reachable on the host/bridge — matching the Pattern-B fallback for internal tools. ## The stack — `stacks/llm/docker-compose.yml` Single service `llama-server`: - **Image:** `ghcr.io/ggml-org/llama.cpp:server-cuda` at a pinned tag. - Infra-pinned per repo convention: label `com.centurylabs.watchtower.enable=false` so Watchtower won't drift it. - **Verification gate:** confirm the pinned prebuilt image includes Pascal `sm_60` kernels and is CUDA ≤ 13 compatible. If it errors on the P100, fall back to a locally built image with `-DCMAKE_CUDA_ARCHITECTURES=60`. - **GPU:** `devices: ["nvidia.com/gpu=0"]` (CDI). - **Volumes:** `/storage1/labdata/llm/models:/models` (bind). - **Command / args (AS DEPLOYED — 64k, required by Hermes' 64K minimum):** - `-m /models/Qwen2.5-14B-Instruct-Q4_K_M.gguf` - `--alias qwen2.5-14b-instruct` (stable model name Hermes references) - `--parallel 1` (one slot gets the FULL context; default 4 slots split it to 32k/seq → fails Hermes) - `-ngl 99` (full offload — 14B Q4 fits in VRAM) - `--ctx-size 65536` - `--rope-scaling yarn --rope-scale 2 --yarn-orig-ctx 32768` (YaRN extends Qwen2.5's 32k native → 64k) - `--override-kv qwen2.context_length=int:65536` (raises GGUF training-context metadata so llama-server does NOT cap the slot back to 32768 — without this the slot is capped and Hermes still sees 32k) - `--flash-attn on` (this build needs the explicit `on` value; a bare `-fa` swallows the next arg) - `--cache-type-k q8_0 --cache-type-v q8_0` (**both q8_0** — q4_0 V-cache is pathological on Pascal: 1.28 tok/s gen at 5–8% GPU util. q8_0/q8_0 → 9.2 tok/s and still fits 64k.) - `--host 0.0.0.0 --port 8080` - API key via `LLAMA_API_KEY` env (env_file) — NOT a CLI flag (no `${VAR}` interpolation; llama-server reads the env var natively) - **Ports:** `"172.20.0.1:8090:8080"` (reachable by Hermes on host + by containers). - **Networks:** private `llm` net only (no `edge` — no Caddy endpoint this round). - **restart:** `unless-stopped`. **Healthcheck:** GET `/health` on 8080. - **`env_file: stack.env`** per repo convention. ### `stacks/llm/stack.env` - `LLM_API_KEY=` (committed per repo policy — secrets are versioned here). ### VRAM budget (Qwen2.5-14B, GQA: 48 layers, 8 KV heads, head_dim 128) - KV cache ≈ 0.375 MiB/token fp16 → **q8_0 halves to ≈ 0.1875 MiB/token**. - Weights Q4_K_M ≈ 9.0 GB; reserve ~0.8 GB compute buffers. - 32k @ q8_0 KV ≈ 6.0 GB → **~15.7 GB total, fits** (tight but safe at 16 GB). ## Model acquisition (one-time) Download `Qwen2.5-14B-Instruct-Q4_K_M.gguf` (~9 GB) from `bartowski/Qwen2.5-14B-Instruct-GGUF` into `/storage1/labdata/llm/models/` on the host (e.g. `huggingface-cli download` or `wget` the single GGUF). Documented host-side step, done before first stack deploy. ## Hermes integration (host-side, not in git) Hermes is the **Nous Research Hermes agent** (`hermes-agent.nousresearch.com`). Its config is `~/.hermes/config.yaml`, which already has a `providers:` list whose entries are exactly OpenAI-compatible upstreams — there's a working `ollama` provider in it today (`type: openai`, `base_url: http://192.168.1.73:11434/v1`). Adding the P100 is one more entry of the same shape; no new integration surface. 1. **Add a provider** to `providers:` in `~/.hermes/config.yaml`: ```yaml - name: valhalla-p100 type: openai base_url: http://172.20.0.1:8090/v1 api_key: models: - qwen2.5-14b-instruct ``` Use `hermes config` / the `hermes` CLI where possible; a direct YAML edit + restart is the fallback (the CLI is the source of truth for `_config_version`). 2. **Select the model** as the active one via `hermes model` (interactive) — or set `model.default: qwen2.5-14b-instruct` (+ matching provider) if it should be the gateway default rather than a switchable option. The current default is `gpt-5.5` / `openai-codex`; we add ours alongside and let the user choose, rather than silently replacing the default. 3. **Restart** the three Hermes processes (gateway dashboard pid-class, `gateway run`, webui) so the new provider/model is live. 4. **Verify** end-to-end: a Hermes prompt routed to `qwen2.5-14b-instruct` produces a completion served by the P100 (confirm via `nvidia-smi` showing the llama-server process holding VRAM during generation). These host-side steps are documented in the plan (and worth a note in CLAUDE.md known-quirks), not committed as repo changes — Hermes isn't in compose, and the alias `qwen2.5-14b-instruct` set via `--alias` is the contract between llama-server and this provider entry. ## Deployment `stacks/llm/` is a new Portainer git stack → must be **registered once** (new stacks aren't auto-created by the poller). Per repo precedent (memory: portainer-new-stack-registration): create via MCP/Portainer, poll `StackList` to confirm, use the stacks' working fine-grained PAT for git creds. Pure `env_file` (empty Portainer UI env). After registration, normal git-push → 5-min poll redeploys apply. ## Final deployed state (verified live 2026-06-26) Portainer stack `llm` (id 34), container `llama-server` healthy. Config: 64k / q8_0 KV / YaRN / `--parallel 1` / `--override-kv qwen2.context_length=int:65536`. VRAM 15.3 GB used, ~0.9 GB free. Hermes `model:` block points at provider `custom` → `http://172.20.0.1:8090/v1` (matched to the `valhalla-p100` entry in the `providers:` list); active model `qwen2.5-14b-instruct`. Original config backed up at `~/.hermes/config.yaml.bak.*`. ### Measured performance (Qwen2.5-14B-Q4_K_M, 64k q8/q8, P100) - **Generation: ~9.2 tok/s** (memory-bound; fine for a personal assistant). - **Prefill: ~54 tok/s** on a large prompt (the misleading ~10 tok/s figure is small-prompt overhead, not throughput). - **Hermes system prompt ≈ 16,400 tokens** → first (cold) turn ≈ **5 min** (all prefill). - **Prompt cache makes it usable:** llama-server matches by longest-common-prefix (`sim_best = 0.999`), so subsequent turns — even new conversations sharing the stable system prompt — reuse the prefix and respond in **~20 s**. The 5 min is a one-time post-restart warmup. ### Hard-won config gotchas (all verified the slow/broken way first) 1. **`-fa` needs an explicit value** in this build: use `--flash-attn on`. A bare `-fa` swallows the next arg (`--cache-type-k`) and crash-loops. 2. **`--parallel 1`** — the default 4 slots split `--ctx-size` to 32k/sequence, which fails Hermes' 64K minimum. One slot serves the full window. 3. **`--override-kv qwen2.context_length=int:65536`** — without it, llama-server *caps the slot back to the GGUF training context (32768)* even with YaRN set, so per-seq stays 32k. 4. **q8_0 V-cache, NOT q4_0** — q4_0 V-cache is pathological on the GP100 (cc 6.0, no DP4A): **1.28 tok/s** generation at 5–8% GPU util. q8_0/q8_0 → 9.2 tok/s and *still* fits 64k. 5. **Hermes requires ≥64K context** and rejects smaller models outright (or set `model.context_length` to override — but then the server must actually serve it). 6. **Hermes provider wiring:** a `providers:` *list* entry is a "named custom provider", activated only by setting the `model:` block to `provider: custom` + matching `base_url`. It is NOT selectable via `--provider ` (that path wants a `providers:` *dict*). ## Non-goals / out of scope - vLLM (ruled out by hardware — see decision above). - SSO/Authentik on the endpoint (LAN/host-only, API-keyed). - A public `llm.ginnoir.com` Caddy endpoint (declined; easy to add later via `edge` + `internal_only`). - Multi-GPU / 2nd P100 install. ## Open risks / follow-ups - **VRAM is tight (~0.9 GB free).** A full 64k prefill held under real load (16k-token Hermes prompt succeeded), but watch for OOM if other GPU users appear; fall back to `--ctx-size 60000` or a smaller weight quant for margin. - **Cold-start latency (~5 min).** Inherent to a 16k system prompt at Pascal prefill speed. Mitigation if it annoys: trim Hermes' prompt (disable `environment_probe`, fewer toolsets) to shrink the cached prefix. - **Prompt-cache persistence across restarts** is in-memory; a container restart re-pays the cold prefill once.