UI: trial hermes-ui + hermes-workspace in parallel, keep winner (new Task 9). Curator stays report-only. eagle-eye trial confirmed (only direct skill-router in the ecosystem; llmtrim noted as complementary). camofox minimal 2-tool skill. All phases now ungated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
255 lines
17 KiB
Markdown
255 lines
17 KiB
Markdown
# Hermes Ecosystem Integration — Triage & Design
|
||
|
||
**Date:** 2026-06-27
|
||
**Author:** Claude (overnight autonomous research for ginnoir; for morning review)
|
||
**Status:** Draft — awaiting ginnoir's decisions on the flagged items below
|
||
|
||
---
|
||
|
||
## 0. TL;DR for the morning
|
||
|
||
You asked me to research seven Hermes-ecosystem repos and plan how each fits. I verified all
|
||
seven are real and active, mapped each against your **actual** Hermes deployment (host systemd
|
||
install on valhalla, llama-swap/P100 backend, ~29-tool curated MCP surface, Codex + Claude Code
|
||
also installed), and triaged them into adopt / trial / decide / skip.
|
||
|
||
**Recommended path:**
|
||
|
||
| Phase | Repo | Verdict | Why |
|
||
|---|---|---|---|
|
||
| **1** | `hermes-agent-acp-skill` | **ADOPT** | Pure skill drop-in; lets the local 20B offload hard work to Codex/Claude Code, which you already run. Highest leverage, lowest risk. |
|
||
| **1** | `hermes-curator-evolver` | **ADOPT (dry-run first)** | Closes Hermes' self-improvement loop; read-only by default; aligns with your curation habit. |
|
||
| **2** | `camofox-browser` | **ADOPT (new homelab stack)** | Gives the local agent real web browsing; clean Docker/Caddy fit; independent of Hermes internals. |
|
||
| **2** | `eagle-eye` | **TRIAL / DEFER** | Solves the exact tool-overload pain you already hit — but very early (5★, 1 commit) and carries a required Chinese-tokenizer dep. Trial behind a flag; don't depend on it. |
|
||
| **3** | `hermes-ui` | **ADOPT (trial)** | Single-HTML + stdlib proxy; stood up next to the bundled webui to evaluate. |
|
||
| **3** | `hermes-workspace` | **ADOPT (trial)** | Full command center; stood up in parallel to compare head-to-head, then keep one. Swarm Mode won't parallelize on one P100 — value is the workspace/observability features. |
|
||
| **3** | `mission-control` | **SKIP** | Not Hermes-native (built for OpenClaw/CrewAI/LangGraph/AutoGen). Redundant with the two Hermes UIs. |
|
||
|
||
**Decisions — RESOLVED 2026-06-27 (see Section 6):** (1) trial **both** UIs, keep the winner;
|
||
(2) curator-evolver stays **report-only** for now; (3) **trial eagle-eye** — it's the only direct
|
||
skill-router in the ecosystem, no more-mature equivalent exists; (4) wire camofox as a **minimal
|
||
2-tool skill**.
|
||
|
||
---
|
||
|
||
## 1. The anchor: your current Hermes deployment
|
||
|
||
Everything below is sized against what you actually run (from project memory, verified against the
|
||
repo and host conventions):
|
||
|
||
- **Hermes Agent v0.17.0 (NousResearch)** runs on the **valhalla host**, not in Docker — systemd
|
||
units `hermes-gateway.service`, `hermes-dashboard.service`, `hermes-webui.service`. Editable
|
||
install at `~/.hermes/hermes-agent/`; config at `~/.hermes/config.yaml`; CLI `~/.local/bin/hermes`.
|
||
Workspace at `/storage1/hermes/workspace/`.
|
||
- **Backend:** llama-swap (`stacks/llm`, Portainer id 34) on a single **Tesla P100 16GB**, serving
|
||
one GGUF at a time, `--parallel 1`, 64k ctx, default `gpt-oss-20b`. OpenAI API at
|
||
`http://172.20.0.1:8090/v1`.
|
||
- **MCP surface deliberately curated to ~29 tools** because gpt-oss-20b drowns in tool overload.
|
||
This is the single most important constraint for everything below: **the local model is weak at
|
||
selection under load, and the GPU runs one model, one sequence at a time.**
|
||
- **You also run Codex (`~/.codex`) and Claude Code**, kept in sync with your Claude global config.
|
||
This is what makes the ACP delegation skill genuinely useful rather than theoretical.
|
||
|
||
### 1.1 Two integration classes (this drives the whole plan)
|
||
|
||
- **Class A — host-side Hermes extensions** (plugins & skills): live in `~/.hermes/` on valhalla,
|
||
applied by SSH + `hermes ...` + gateway restart. **These are NOT homelabstack repo content.**
|
||
They are host-managed exactly like the rustdesk/obsidian/xvfb units — tracked in project memory
|
||
and the Obsidian vault, not committed here. Fully reversible (config backups + plugin uninstall).
|
||
→ `hermes-agent-acp-skill`, `hermes-curator-evolver`, `eagle-eye`.
|
||
- **Class B — Docker services**: become `stacks/<name>/` entries, deployed via the normal
|
||
Gitea-poll path, fronted by Caddy (`internal_only` + Authentik), secrets in `stack.env`,
|
||
infra images pinned out of Watchtower.
|
||
→ `camofox-browser`, and (if adopted) `hermes-ui` / `hermes-workspace` / `mission-control`.
|
||
|
||
---
|
||
|
||
## 2. Per-repo analysis
|
||
|
||
### 2.1 `hermes-agent-acp-skill` (Rainhoole) — **ADOPT, Phase 1**
|
||
|
||
- **What:** A Hermes *skill* (`SKILL.md` + `README.md`, MIT) implementing `delegate_task()` —
|
||
multi-agent delegation across Hermes internal subagents, **Codex**, and **Claude Code**. Handles
|
||
agent routing (`agent=...`), context isolation, and safety controls (recommended 900 s external
|
||
timeout, 24,000-char output cap).
|
||
- **Fit:** Strong and specific. Your local model is gpt-oss-20b — fine for orchestration and light
|
||
tasks, weak for real implementation/review. This skill lets Hermes *route the hard parts to the
|
||
frontier CLIs you already have installed and configured*. It turns the weak local brain into a
|
||
dispatcher.
|
||
- **Risk:** Low. It's a skill file, not running code you didn't write — but it *invokes external
|
||
agents*, so the real work is verifying Codex/Claude Code are launchable from the Hermes service
|
||
environment (PATH, auth, non-interactive flags) and that timeouts/output caps are sane.
|
||
- **Implementation sketch:** Drop into `~/.hermes` skills dir (or `hermes skills install` if the
|
||
repo packages it that way — verify on the host); set delegation targets to the Codex + Claude
|
||
Code invocations; restart gateway; smoke-test a trivial delegated task end-to-end.
|
||
|
||
### 2.2 `hermes-curator-evolver` (pingchesu) — **ADOPT, dry-run first, Phase 1**
|
||
|
||
- **What:** Local-first Hermes *plugin* (Python 3.11, SQLite, `hermes plugins install`). Mines
|
||
evidence from sessions/tool-calls, ranks skill-improvement candidates, and generates **dry-run,
|
||
reviewable** proposals. **Read-only by default; mutations require explicit `--approve`.** Only
|
||
touches *local agent-created* skills — skips bundled, hub-installed, and external skills. Size
|
||
guardrails, backup/rollback manifests, optional Qwen-embeddings + BGE-reranker for semantic
|
||
ranking, native scheduler (systemd) integration.
|
||
- **Fit:** This *is* Hermes' "self-improving" thesis made operational, and it matches how you
|
||
already work (curate, gate, keep things reversible). The default posture (report-only, dry-run,
|
||
human-review queue) is exactly the safe on-ramp.
|
||
- **Risk / caveats:**
|
||
- It writes to skills with `--approve`. **Recommendation: run report-only / dry-run indefinitely
|
||
at first**; promote to `--approve` only after you've seen a few proposal batches you'd have
|
||
accepted anyway.
|
||
- Optional embeddings/reranker want a model. **Do not load them onto the P100** alongside
|
||
gpt-oss — the GPU is single-model/`--parallel 1`. Run ranking on **CPU**, or skip the dense
|
||
layer initially (BM25/FTS ranking is enough to start).
|
||
- Schedule the mining pass **off-peak** via a systemd timer so it never contends with live
|
||
inference.
|
||
- **Implementation sketch:** `hermes plugins install` on the host; point it at the session store;
|
||
run one manual `collect`→`rank`→`report` cycle by hand; inspect the review queue; only then wire
|
||
a nightly systemd timer in report-only mode.
|
||
|
||
### 2.3 `eagle-eye` (willingning-coder) — **TRIAL / DEFER, Phase 2**
|
||
|
||
- **What:** Zero-invasive skill *pre-filter* plugin. Before the LLM call, narrows 50+ skills to the
|
||
top ~5 via a 5-layer cascade (hard triggers → FTS5/BM25 → synonyms → dense embeddings → RRF
|
||
fusion), degrading gracefully when optional deps are absent. "Not matching is a valid result."
|
||
- **Fit (conceptual):** Excellent — this is the **skill-side analog of the MCP tool-overload fix
|
||
you already did**. Fewer, better skill candidates in the prompt = better selection by a weak
|
||
model + less Pascal prefill. If it works, it directly improves gpt-oss-20b's behavior.
|
||
- **Risk / caveats (why it's a trial, not an adopt):**
|
||
- **Maturity:** 5 stars, single commit. Early. Treat as experimental.
|
||
- **Dependency smell:** `jieba` (Chinese tokenization) is a *required* dep; your skill library is
|
||
English. Works, but signals the author's primary use case differs from yours — watch for
|
||
English-tokenization edge cases and maintenance drift.
|
||
- Same P100 rule: keep its dense-embedding layer on **CPU** or disabled; lean on L1–L2
|
||
(deterministic + BM25) first.
|
||
- **Implementation sketch:** Install behind an easy on/off switch on the host; A/B a handful of
|
||
representative prompts with it on vs off; keep it **only if** measured skill-selection quality
|
||
improves. Do not let any workflow depend on it.
|
||
|
||
### 2.4 `camofox-browser` (jo-inc) — **ADOPT as a homelab stack, Phase 2**
|
||
|
||
- **What:** Stealth headless browser (Camoufox/Firefox) exposing a **REST API for AI agents** —
|
||
anti-fingerprinting, element-reference interaction (`e1`,`e2`), session isolation, proxy/GeoIP,
|
||
YouTube transcripts, Docker/Fly/Railway deploy. 7.2k★, actively maintained. **No Hermes coupling**
|
||
— it's a generic tool backend.
|
||
- **Fit:** Clean Class-B homelab fit and a real capability gap-filler: your agent currently has
|
||
freshrss + obsidian but no general web browser. As `stacks/camofox`, internal-only behind Caddy,
|
||
it gives Hermes (and Codex/Claude) reliable browsing.
|
||
- **Risk / caveats:**
|
||
- It's anti-detection tooling — appropriate for personal research; **keep it strictly LAN/SSO-gated
|
||
(`import internal_only` + Authentik)** and off the public internet. New attack surface; pin the
|
||
image and read the Dockerfile before running.
|
||
- Heavier service (~300 MB Camoufox binary, Firefox). Fine on valhalla, but it's not free.
|
||
- **Wiring to Hermes:** either a thin MCP shim around its REST API, or a Hermes skill that curls
|
||
it. Given your tool-overload sensitivity, expose it as **one or two** high-level tools (e.g.
|
||
`browse(url)`, `search(query)`), not its full surface.
|
||
- **Implementation sketch:** New `stacks/camofox/docker-compose.yml` (bind config to `/config/camofox`,
|
||
blobs/cache to `/storage1/labdata/camofox`, join `edge`), `stack.env`, a Caddy block
|
||
(`camofox.ginnoir.com`, `import internal_only`), regenerate bookmarks, push. Then add the minimal
|
||
Hermes tool/skill wrapper on the host and restart the gateway.
|
||
|
||
### 2.5 The three dashboards — pick **at most one** (Phase 3, DECISION)
|
||
|
||
You already run the bundled `hermes-webui.service`, so a UI is an *upgrade*, not a gap. These three
|
||
overlap heavily; adopting more than one is waste.
|
||
|
||
| | `hermes-ui` (pyrate-llama) | `hermes-workspace` (outsourc-e) | `mission-control` (builderz-labs) |
|
||
|---|---|---|---|
|
||
| Stars | 173 | 5.9k | 5.4k |
|
||
| Stack | 1 HTML file + Python **stdlib** proxy (`serve_lite.py`, :3333 → gateway :8642) | React/TS/**Vite build**, Node 22+, → gateway :8642 + dashboard :9119 | Next.js 16 / React 19, better-sqlite3 |
|
||
| Hermes-native? | **Yes** | **Yes** (best with vanilla hermes-agent) | **No** (OpenClaw/CrewAI/LangGraph/AutoGen/Claude SDK) |
|
||
| Effort to host | **Trivial** (no build) | Medium (build → container/stack) | Medium + **adapter work** |
|
||
| Headline features | chat, tasks/kanban, files, terminal, skills, MCP browser, cron, memory, health | all of the above + ops dashboards, "Conductor" missions, **Swarm Mode (tmux worker pools)** | 32 panels, RBAC, spend/trust scoring, multi-framework |
|
||
| Verdict | **Trial first** | **Consider as a proper stack later** | **Skip / defer** |
|
||
|
||
- **`hermes-ui` — recommended first move.** Lowest cost to evaluate a richer UI. Can run as a tiny
|
||
host service or a minimal container next to the bundled webui; no build pipeline. If it's better
|
||
than the bundled webui, keep it; if not, delete one file.
|
||
- **`hermes-workspace` — the ambitious option.** Genuinely the most capable, best run as a real
|
||
`stacks/hermes-workspace` Docker build behind Caddy + Authentik. **Caveat to flag:** its
|
||
marquee **Swarm Mode (parallel tmux worker pools) cannot actually parallelize inference on your
|
||
setup** — the single P100 runs one model with `--parallel 1`, so "swarm" workers serialize on the
|
||
GPU (or must use a different/remote backend per worker). Adopt for the workspace/observability
|
||
features, not for swarm, unless/until you add GPU capacity or point swarm workers at cloud
|
||
providers.
|
||
- **`mission-control` — skip unless requirements change.** It's framework-agnostic and **not
|
||
Hermes-native** — using it means writing a Hermes adapter for a dashboard whose job the two
|
||
Hermes-native UIs already do. Only worth it if you later want a single ops plane across *multiple*
|
||
agent frameworks (you don't run those today).
|
||
|
||
---
|
||
|
||
## 3. Recommended phased rollout
|
||
|
||
**Phase 1 — host-side, reversible, high-leverage (no repo changes):**
|
||
1. `hermes-agent-acp-skill` — install skill, wire Codex + Claude Code delegation targets, smoke-test.
|
||
2. `hermes-curator-evolver` — install plugin in **report-only/dry-run**, manual cycle, then nightly
|
||
systemd timer (report-only). CPU ranking; no P100 embeddings.
|
||
|
||
**Phase 2 — new capability + experiment:**
|
||
3. `camofox-browser` — new `stacks/camofox` (Class B), Caddy internal-only + Authentik, minimal
|
||
Hermes tool wrapper.
|
||
4. `eagle-eye` — host-side trial behind an on/off switch; A/B skill-selection; keep only if it helps.
|
||
|
||
**Phase 3 — UI decision (your call):**
|
||
5. Trial `hermes-ui` alongside the bundled webui. Evaluate `hermes-workspace` as a stack if the
|
||
workspace features justify the build. Skip `mission-control`.
|
||
|
||
Each phase is independently shippable and independently reversible. Stop after any phase.
|
||
|
||
---
|
||
|
||
## 4. Cross-cutting concerns (apply to every adoption)
|
||
|
||
- **Single-P100 discipline.** One GPU, one model, `--parallel 1`. Anything implying parallel or
|
||
extra model loads — curator embeddings, eagle-eye dense layer, workspace swarm — must run on
|
||
**CPU**, **off-peak**, or **not at all**. Never co-resident with gpt-oss on the P100.
|
||
- **Host vs repo boundary.** Class-A extensions are host-managed in `~/.hermes/` (document in memory
|
||
+ vault, like the rustdesk/obsidian units). Only Class-B services touch this repo. Don't
|
||
accidentally commit host config here.
|
||
- **Tool/skill surface minimalism.** The whole reason your MCP is curated to ~29 tools. Expose new
|
||
capabilities (camofox, delegation) as a *few high-level* tools, not raw surfaces, or you reintroduce
|
||
the overload you fixed.
|
||
- **Provenance & safety.** Several repos are small/young (eagle-eye 1 commit; acp-skill 1 commit).
|
||
**Read the code before running it**, especially anything that writes skills (curator) or executes
|
||
shell / external agents (acp-skill, camofox). **Pin to a reviewed commit/tag**; keep stack images
|
||
out of Watchtower auto-update.
|
||
- **Reversibility.** Back up `~/.hermes/config.yaml` before each Class-A change (you already do this);
|
||
snapshot before promoting curator to `--approve`.
|
||
- **Security posture.** Class-B services: `import internal_only` + Authentik, secrets in `stack.env`
|
||
(LF endings), never public.
|
||
|
||
---
|
||
|
||
## 5. What I am explicitly NOT proposing
|
||
|
||
- Not running curator-evolver in `--approve` (auto-skill-editing) mode out of the gate.
|
||
- Not loading any second model onto the P100 (no GPU embeddings, no GPU swarm).
|
||
- Not adopting more than one dashboard.
|
||
- Not committing host-side Hermes plugins/skills into the homelabstack repo.
|
||
- Not exposing camofox or any new UI to the public internet.
|
||
|
||
---
|
||
|
||
## 6. DECISIONS — RESOLVED 2026-06-27
|
||
|
||
1. **UI direction → trial BOTH.** Stand up `hermes-ui` *and* `hermes-workspace` in parallel,
|
||
evaluate head-to-head against the bundled webui, then keep the winner and remove the other.
|
||
(Phase 3 now builds both.)
|
||
2. **Curator autonomy → report-only for now.** `hermes-curator-evolver` runs dry-run/report only;
|
||
no `--approve` / `--apply-low-risk` / `--approve-auto-apply`. Revisit after a trust period.
|
||
3. **eagle-eye → trial it.** Rationale (ginnoir): wants a large skill library available without
|
||
weighting every turn — which is exactly eagle-eye's pre-LLM skill-selection job. **Research
|
||
result:** eagle-eye is the *only* direct skill-router in the Hermes ecosystem (per
|
||
`awesome-hermes-agent` / Hermes Atlas); the nominal "alternatives" do different jobs —
|
||
`llmtrim` (compresses tool schemas + MCP output; **complementary**, worth a later look for the
|
||
tool side), `rtk-hermes` (shell-output compression), `hermes-snow-search` (faster session
|
||
search), `hermes-motif` (mines traces into micro-skills; a curator-evolver *rival*, not a
|
||
router). So eagle-eye stands, with the trial caveat (1-commit maturity, `jieba` dep): keep only
|
||
if it measurably improves selection; run L1+L2 only (no P100 embeddings).
|
||
4. **camofox wiring → minimal.** A small 2-tool Hermes skill (`browse`, `search`) over the REST
|
||
API, not the full surface or an MCP shim — protects the curated tool budget.
|
||
|
||
All four resolved ⇒ **Phases 1, 2, and 3 are all actionable.** The implementation plan that
|
||
follows details every phase.
|