ginnoirandClaude Opus 4.8 74791105b3 fix(llm): --parallel 1 so a single request gets the full 64k context
With the default 4 slots, llama-server splits ctx into 32k per sequence, which
fails Hermes' 64K minimum. One slot serves the full 65536 per request (serial
agent use; concurrent calls queue).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:42:41 -05:00
2026-06-14 22:53:22 -05:00
2026-06-02 21:32:30 -05:00
2026-06-02 21:32:30 -05:00
S
Description
Mirror of ginnoir/homelabstack (primary on GitHub).
Readme MIT
753 KiB
Languages
Python 54.5%
Shell 18.5%
PowerShell 16%
JavaScript 5.5%
HTML 4.7%
Other 0.7%