feat(llm): llama-swap multi-model menu (gpt-oss-20b default + gemma4 family)
Replace single llama-server with llama-swap so all benchmarked models are selectable from Hermes' menu and hot-swapped on the one P100. Menu: gpt-oss-20b (default, ~45s cold start), gemma-4-26b-a4b (MoE), gemma-4-12b, gemma-4-e4b. qwen3-30b-a3b excluded (OOMs at 64k in 16GB). All 64k, q8/q8 KV, --parallel 1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
eba51a52a7
commit
63fd3fd1cc
@@ -1,61 +1,31 @@
|
||||
# llm stack — local LLM inference backend for the Hermes agent.
|
||||
# llm stack — model-swapping LLM backend for the Hermes agent.
|
||||
#
|
||||
# Single service: llama.cpp's OpenAI-compatible server (llama-server) serving
|
||||
# Qwen2.5-14B-Instruct (Q4_K_M GGUF) on the host's Tesla P100-16GB via CDI.
|
||||
# Chosen over vLLM because the P100 (GP100, compute capability 6.0) lacks the
|
||||
# DP4A INT8 instructions vLLM's AWQ/GPTQ kernels require — see
|
||||
# docs/superpowers/specs/2026-06-26-llm-backend-hermes-design.md.
|
||||
# llama-swap fronts multiple GGUF models on the single Tesla P100 (16GB). Only one
|
||||
# model fits in VRAM at a time, so llama-swap presents all of them via /v1/models
|
||||
# and hot-swaps on demand (selecting a different model = a few-second reload). The
|
||||
# per-model llama-server commands + args live in llama-swap-config.yaml.
|
||||
#
|
||||
# Pure env_file (LLAMA_API_KEY) — no Portainer UI env, no ${VAR} interpolation.
|
||||
# Image is infra-pinned out of Watchtower (manual tag bumps only).
|
||||
# The bundled llama.cpp in llama-swap:cuda is build 9803 (5c7c22c3e) — the same
|
||||
# build validated on this Pascal card for gemma4 + gpt-oss. Default model and the
|
||||
# selectable menu are driven from Hermes (~/.hermes/config.yaml: model.default =
|
||||
# gpt-oss-20b; provider valhalla-p100 models: list = the keys in the swap config).
|
||||
#
|
||||
# The server's OpenAI API is published on the host at 172.20.0.1:8090 (the edge
|
||||
# bridge gateway, a local host IP). Host-side Hermes reaches it there directly;
|
||||
# no Caddy block this round. Model weights live on the ZFS tier; the
|
||||
# /storage1/labdata/llm/models dir is pre-created with the GGUF before deploy.
|
||||
# Endpoint published on 172.20.0.1:8090 (edge bridge gateway, a host IP) for the
|
||||
# host-side Hermes agent. Internal-only; no Caddy, no auth (LAN/host-only).
|
||||
# Image is infra-pinned out of Watchtower.
|
||||
services:
|
||||
llama-server:
|
||||
image: ghcr.io/ggml-org/llama.cpp:server-cuda
|
||||
container_name: llama-server
|
||||
llama-swap:
|
||||
image: ghcr.io/mostlygeek/llama-swap:cuda
|
||||
container_name: llama-swap
|
||||
restart: unless-stopped
|
||||
labels:
|
||||
- "com.centurylabs.watchtower.enable=false"
|
||||
networks: [llm]
|
||||
env_file:
|
||||
- stack.env
|
||||
devices:
|
||||
- "nvidia.com/gpu=0"
|
||||
volumes:
|
||||
- /storage1/labdata/llm/models:/models
|
||||
command:
|
||||
- "-m"
|
||||
- "/models/Qwen2.5-14B-Instruct-Q4_K_M.gguf"
|
||||
- "--alias"
|
||||
- "qwen2.5-14b-instruct"
|
||||
- "--parallel"
|
||||
- "1"
|
||||
- "-ngl"
|
||||
- "99"
|
||||
- "--ctx-size"
|
||||
- "65536"
|
||||
- "--rope-scaling"
|
||||
- "yarn"
|
||||
- "--rope-scale"
|
||||
- "2"
|
||||
- "--yarn-orig-ctx"
|
||||
- "32768"
|
||||
- "--override-kv"
|
||||
- "qwen2.context_length=int:65536"
|
||||
- "--flash-attn"
|
||||
- "on"
|
||||
- "--cache-type-k"
|
||||
- "q8_0"
|
||||
- "--cache-type-v"
|
||||
- "q8_0"
|
||||
- "--host"
|
||||
- "0.0.0.0"
|
||||
- "--port"
|
||||
- "8080"
|
||||
- ./llama-swap-config.yaml:/app/config.yaml:ro
|
||||
ports:
|
||||
- "172.20.0.1:8090:8080"
|
||||
healthcheck:
|
||||
@@ -63,7 +33,7 @@ services:
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 5
|
||||
start_period: 180s
|
||||
start_period: 30s
|
||||
|
||||
networks:
|
||||
llm:
|
||||
|
||||
Reference in New Issue
Block a user