74791105b3927d5d1883cbcf424176e5dd70c1fe
With the default 4 slots, llama-server splits ctx into 32k per sequence, which fails Hermes' 64K minimum. One slot serves the full 65536 per request (serial agent use; concurrent calls queue). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
homelabstack
Description
Mirror of ginnoir/homelabstack (primary on GitHub).
824 KiB
0 Stars
1 Watchers
0 Forks
Languages
Python
54.3%
Shell
18.5%
PowerShell
15.9%
JavaScript
5.5%
HTML
4.6%
Other
1.1%