Initial commit: llm-router — smart OpenAI/Anthropic/Ollama router

A stdlib pre-router in front of Ollama with LiteLLM backend:
- auto model selection by content/tools/modality, with fallbacks
- OpenAI /v1, Anthropic /v1/messages, and Ollama-native /api/* endpoints
- Whisper-shaped /v1/audio/transcriptions + in-chat audio
- key-based fleet policies (e.g. force a client onto uncensored models)
- optional Bearer auth; launchd/systemd service install
- benchmark harnesses (speed, quality, agentic tool use) with sample results

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Joseph Costa
2026-07-05 02:05:16 -05:00
co-authored by Claude Opus 4.8
commit 9938d46a67
32 changed files with 2419 additions and 0 deletions
+9
View File
@@ -0,0 +1,9 @@
Copy this file to `.uncensored_key` and put a single secret token in it.
Any request whose `Authorization: Bearer <token>` matches this value is FORCED
onto the uncensored model fleet, regardless of the model it asks for. Use it to
pin a specific client (e.g. a family assistant app) to uncensored models.
This is independent of `.apikey` — the router can be open (no auth) and still
honor this key to switch a client to the uncensored fleet.
Generate one: echo "famapp-$(openssl rand -hex 16)"