A stdlib pre-router in front of Ollama with LiteLLM backend: - auto model selection by content/tools/modality, with fallbacks - OpenAI /v1, Anthropic /v1/messages, and Ollama-native /api/* endpoints - Whisper-shaped /v1/audio/transcriptions + in-chat audio - key-based fleet policies (e.g. force a client onto uncensored models) - optional Bearer auth; launchd/systemd service install - benchmark harnesses (speed, quality, agentic tool use) with sample results Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
19 lines
237 B
Plaintext
19 lines
237 B
Plaintext
# secrets — never commit real keys
|
|
.apikey
|
|
.uncensored_key
|
|
|
|
# python / venv
|
|
.venv/
|
|
__pycache__/
|
|
*.pyc
|
|
|
|
# runtime state
|
|
*.pid
|
|
*.log
|
|
|
|
# scratch model definitions (generated by scripts/make-context-variant.sh)
|
|
Modelfile.*
|
|
|
|
# OS
|
|
.DS_Store
|