13424fcf75875669169a6e2e91e1d2b2177160d4
llama-server caps the slot to the GGUF training context (32768) and ignores the YaRN-extended size, leaving per-seq context at 32k. Raise qwen2.context_length metadata to 65536 so the full window is served per request. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
homelabstack
Description
Mirror of ginnoir/homelabstack (primary on GitHub).
824 KiB
0 Stars
1 Watchers
0 Forks
Languages
Python
54.3%
Shell
18.5%
PowerShell
15.9%
JavaScript
5.5%
HTML
4.6%
Other
1.1%