10 Commits
Author SHA1 Message Date
ginnoirandClaude Opus 4.8 009a474e90 feat(llm): add ornith-1.0-9b coding model to llama-swap menu
DeepReinforce Ornith-1.0 (dense 9B on Qwen 3.5, Q5_K_M, MIT), an
agentic-coding model. Tool-calls + <think> work under --jinja; native
256k so no YaRN. Loads at ~7.7GB VRAM @ 64k.

Benchmark (docs/2026-06-27-ornith-9b-benchmark.md): quality ties
gpt-oss-20b but gen is ~2.5-3x slower (dense 9B active vs gpt-oss MoE
3.6B active on the compute-bound P100). Default stays gpt-oss-20b;
ornith kept as a coding specialist in the menu.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 13:51:19 -05:00
ginnoirandClaude Opus 4.8 6cef600d25 fix(llm): add --jinja so gpt-oss harmony template returns content
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 00:09:09 -05:00
ginnoirandClaude Opus 4.8 f20712a5d8 fix(llm): bind llama-swap config from absolute /config/llm (Portainer rel-bind)
Portainer's git checkout auto-creates a relative repo-file bind as a directory,
breaking the /app/config.yaml mount. Use the absolute host path like the share
stack; repo copy stays canonical, mirrored to /config/llm on deploy.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 00:04:46 -05:00
ginnoirandClaude Opus 4.8 63fd3fd1cc feat(llm): llama-swap multi-model menu (gpt-oss-20b default + gemma4 family)
Replace single llama-server with llama-swap so all benchmarked models are
selectable from Hermes' menu and hot-swapped on the one P100. Menu: gpt-oss-20b
(default, ~45s cold start), gemma-4-26b-a4b (MoE), gemma-4-12b, gemma-4-e4b.
qwen3-30b-a3b excluded (OOMs at 64k in 16GB). All 64k, q8/q8 KV, --parallel 1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 00:02:58 -05:00
ginnoirandClaude Opus 4.8 4aa2cd9468 fix(llm): q8_0 V-cache (q4_0 tanked generation to 1.3 tok/s on Pascal)
The q4_0 V-cache + flash-attention path is pathological on the GP100: 1.28
tok/s generation at 5-8% GPU util. q8_0 V-cache gives 9.2 tok/s and still fits
64k context in 16GB (15.3GB used, ~950MB free).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 17:04:17 -05:00
ginnoirandClaude Opus 4.8 13424fcf75 fix(llm): override-kv context_length=65536 so slot isn't capped to 32k
llama-server caps the slot to the GGUF training context (32768) and ignores the
YaRN-extended size, leaving per-seq context at 32k. Raise qwen2.context_length
metadata to 65536 so the full window is served per request.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:45:33 -05:00
ginnoirandClaude Opus 4.8 74791105b3 fix(llm): --parallel 1 so a single request gets the full 64k context
With the default 4 slots, llama-server splits ctx into 32k per sequence, which
fails Hermes' 64K minimum. One slot serves the full 65536 per request (serial
agent use; concurrent calls queue).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:42:41 -05:00
ginnoirandClaude Opus 4.8 265b407e8d feat(llm): serve 64k context (YaRN) to meet Hermes' 64K minimum
Hermes Agent rejects models with <64K context. Qwen2.5-14B is 32k native, so
enable YaRN rope-scaling (2x → 65536) and drop the V-cache to q4_0 for VRAM
headroom on the 16GB P100.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:39:56 -05:00
ginnoirandClaude Opus 4.8 605c6d3709 fix(llm): use --flash-attn on (this llama.cpp build requires explicit value)
The server-cuda image parses -fa as --flash-attn [on|off|auto], so a bare -fa
swallowed the following --cache-type-k as its value and crash-looped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:31:01 -05:00
ginnoirandClaude Opus 4.8 8e7682985d feat(llm): add llama.cpp inference stack for Hermes (Qwen2.5-14B on P100)
New stacks/llm/ serves Qwen2.5-14B-Instruct (Q4_K_M GGUF) via llama.cpp's
OpenAI-compatible server on the Tesla P100 (CDI nvidia.com/gpu=0), published on
172.20.0.1:8090 for the host-side Hermes agent. vLLM was rejected: the P100
(cc 6.0) lacks the DP4A INT8 instructions its AWQ/GPTQ kernels need.

Includes design spec and implementation plan under docs/superpowers/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:22:57 -05:00