63fd3fd1ccdce98282b7eb031af6a0ef782dffe0
Replace single llama-server with llama-swap so all benchmarked models are selectable from Hermes' menu and hot-swapped on the one P100. Menu: gpt-oss-20b (default, ~45s cold start), gemma-4-26b-a4b (MoE), gemma-4-12b, gemma-4-e4b. qwen3-30b-a3b excluded (OOMs at 64k in 16GB). All 64k, q8/q8 KV, --parallel 1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
homelabstack
Languages
Python
54.5%
Shell
18.5%
PowerShell
16%
JavaScript
5.5%
HTML
4.7%
Other
0.7%