ginnoirandClaude Opus 4.8 13424fcf75 fix(llm): override-kv context_length=65536 so slot isn't capped to 32k
llama-server caps the slot to the GGUF training context (32768) and ignores the
YaRN-extended size, leaving per-seq context at 32k. Raise qwen2.context_length
metadata to 65536 so the full window is served per request.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:45:33 -05:00
2026-06-14 22:53:22 -05:00
2026-06-02 21:32:30 -05:00
2026-06-02 21:32:30 -05:00
S
Description
Mirror of ginnoir/homelabstack (primary on GitHub).
Readme MIT
824 KiB
0 Stars 1 Watchers 0 Forks
Languages
Python 54.3%
Shell 18.5%
PowerShell 15.9%
JavaScript 5.5%
HTML 4.6%
Other 1.1%