4aa2cd9468c4bde379d9c10c79b418d9ce0a1f03
The q4_0 V-cache + flash-attention path is pathological on the GP100: 1.28 tok/s generation at 5-8% GPU util. q8_0 V-cache gives 9.2 tok/s and still fits 64k context in 16GB (15.3GB used, ~950MB free). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
homelabstack
Languages
Python
54.5%
Shell
18.5%
PowerShell
16%
JavaScript
5.5%
HTML
4.7%
Other
0.7%