llama-server caps the slot to the GGUF training context (32768) and ignores the
YaRN-extended size, leaving per-seq context at 32k. Raise qwen2.context_length
metadata to 65536 so the full window is served per request.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>