ginnoirandClaude Opus 4.8 8e7682985d feat(llm): add llama.cpp inference stack for Hermes (Qwen2.5-14B on P100)
New stacks/llm/ serves Qwen2.5-14B-Instruct (Q4_K_M GGUF) via llama.cpp's
OpenAI-compatible server on the Tesla P100 (CDI nvidia.com/gpu=0), published on
172.20.0.1:8090 for the host-side Hermes agent. vLLM was rejected: the P100
(cc 6.0) lacks the DP4A INT8 instructions its AWQ/GPTQ kernels need.

Includes design spec and implementation plan under docs/superpowers/.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:22:57 -05:00
2026-06-14 22:53:22 -05:00
2026-06-02 21:32:30 -05:00
2026-06-02 21:32:30 -05:00
S
Description
Mirror of ginnoir/homelabstack (primary on GitHub).
Readme MIT
753 KiB
Languages
Python 54.5%
Shell 18.5%
PowerShell 16%
JavaScript 5.5%
HTML 4.7%
Other 0.7%