llama-server caps the slot to the GGUF training context (32768) and ignores the
YaRN-extended size, leaving per-seq context at 32k. Raise qwen2.context_length
metadata to 65536 so the full window is served per request.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
With the default 4 slots, llama-server splits ctx into 32k per sequence, which
fails Hermes' 64K minimum. One slot serves the full 65536 per request (serial
agent use; concurrent calls queue).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hermes Agent rejects models with <64K context. Qwen2.5-14B is 32k native, so
enable YaRN rope-scaling (2x → 65536) and drop the V-cache to q4_0 for VRAM
headroom on the 16GB P100.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The server-cuda image parses -fa as --flash-attn [on|off|auto], so a bare -fa
swallowed the following --cache-type-k as its value and crash-looped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New stacks/llm/ serves Qwen2.5-14B-Instruct (Q4_K_M GGUF) via llama.cpp's
OpenAI-compatible server on the Tesla P100 (CDI nvidia.com/gpu=0), published on
172.20.0.1:8090 for the host-side Hermes agent. vLLM was rejected: the P100
(cc 6.0) lacks the DP4A INT8 instructions its AWQ/GPTQ kernels need.
Includes design spec and implementation plan under docs/superpowers/.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>