Hermes Agent rejects models with <64K context. Qwen2.5-14B is 32k native, so enable YaRN rope-scaling (2x → 65536) and drop the V-cache to q4_0 for VRAM headroom on the 16GB P100. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>