A stdlib pre-router in front of Ollama with LiteLLM backend: - auto model selection by content/tools/modality, with fallbacks - OpenAI /v1, Anthropic /v1/messages, and Ollama-native /api/* endpoints - Whisper-shaped /v1/audio/transcriptions + in-chat audio - key-based fleet policies (e.g. force a client onto uncensored models) - optional Bearer auth; launchd/systemd service install - benchmark harnesses (speed, quality, agentic tool use) with sample results Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2.1 KiB
2.1 KiB
Benchmarks
Four self-contained harnesses in benchmarks/ score your local
fleet on speed, quality, and tool use. They talk to Ollama (:11434) and the
router (:8080) directly. Run them with the venv Python, from the benchmarks/
directory (they import each other):
cd benchmarks
../.venv/bin/python benchmark.py # speed + OCR + routing correctness
../.venv/bin/python quality_bench.py # objective quality (code exec, math, MC, instr)
../.venv/bin/python agent_bench.py # basic tool-calling
../.venv/bin/python hard_agent_bench.py # multi-turn agentic (seq/parallel/recover/abstain)
Each writes a *_report.md (and some a *_results.json). Sample outputs from the
reference machine (Apple M4 Max, 36 GB) are checked in under benchmarks/results/.
What each measures
- benchmark.py — generation & prompt tok/s and GPU placement (from Ollama's own eval metrics), an OCR pass on vision models, and live routing-decision checks.
- quality_bench.py — objective auto-grading: executed unit tests (coding), exact-match math (GSM8K-style), multiple-choice, and programmatic instruction-following. Runs with thinking disabled for a fast, comparable baseline.
- agent_bench.py — tool-calling basics via Ollama's tools API: invoke, args, selection, use-result, and abstain (not calling a tool when none is needed).
- hard_agent_bench.py — a real multi-turn tool loop (tools actually execute and
results feed back): sequential dependency, parallel/multi-entity, 3-hop chains,
tool selection under distractors, error recovery, and no-fabrication honesty.
Pass model names as args to limit the set:
... hard_agent_bench.py glm-4.7-flash gemma4:26b.
Notes
- These assume the same model names as the reference fleet — edit the model lists at the top of each script to match yours.
quality_bench.py/benchmark.pygenerate a test image with Pillow (in the venv) for the OCR/vision checks.- Objective grading (running generated code, exact-match answers) is deliberately used over LLM-as-judge to avoid judge bias.