13424fcf75875669169a6e2e91e1d2b2177160d4
llama-server caps the slot to the GGUF training context (32768) and ignores the YaRN-extended size, leaving per-seq context at 32k. Raise qwen2.context_length metadata to 65536 so the full window is served per request. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
homelabstack
Languages
Python
54.5%
Shell
18.5%
PowerShell
16%
JavaScript
5.5%
HTML
4.7%
Other
0.7%