8e7682985dd2073755d2ed608e425e944a53f0d4
New stacks/llm/ serves Qwen2.5-14B-Instruct (Q4_K_M GGUF) via llama.cpp's OpenAI-compatible server on the Tesla P100 (CDI nvidia.com/gpu=0), published on 172.20.0.1:8090 for the host-side Hermes agent. vLLM was rejected: the P100 (cc 6.0) lacks the DP4A INT8 instructions its AWQ/GPTQ kernels need. Includes design spec and implementation plan under docs/superpowers/. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
homelabstack
Languages
Python
54.5%
Shell
18.5%
PowerShell
16%
JavaScript
5.5%
HTML
4.7%
Other
0.7%