# Benchmarks Four self-contained harnesses in [`benchmarks/`](../benchmarks/) score your local fleet on speed, quality, and tool use. They talk to Ollama (`:11434`) and the router (`:8080`) directly. Run them with the venv Python, from the `benchmarks/` directory (they import each other): ```bash cd benchmarks ../.venv/bin/python benchmark.py # speed + OCR + routing correctness ../.venv/bin/python quality_bench.py # objective quality (code exec, math, MC, instr) ../.venv/bin/python agent_bench.py # basic tool-calling ../.venv/bin/python hard_agent_bench.py # multi-turn agentic (seq/parallel/recover/abstain) ``` Each writes a `*_report.md` (and some a `*_results.json`). Sample outputs from the reference machine (Apple M4 Max, 36 GB) are checked in under `benchmarks/results/`. ## What each measures - **benchmark.py** — generation & prompt tok/s and GPU placement (from Ollama's own eval metrics), an OCR pass on vision models, and live routing-decision checks. - **quality_bench.py** — objective auto-grading: executed unit tests (coding), exact-match math (GSM8K-style), multiple-choice, and programmatic instruction-following. Runs with thinking disabled for a fast, comparable baseline. - **agent_bench.py** — tool-calling basics via Ollama's tools API: invoke, args, selection, use-result, and *abstain* (not calling a tool when none is needed). - **hard_agent_bench.py** — a real multi-turn tool loop (tools actually execute and results feed back): sequential dependency, parallel/multi-entity, 3-hop chains, tool selection under distractors, error recovery, and no-fabrication honesty. Pass model names as args to limit the set: `... hard_agent_bench.py glm-4.7-flash gemma4:26b`. ## Notes - These assume the same model names as the reference fleet — edit the model lists at the top of each script to match yours. - `quality_bench.py` / `benchmark.py` generate a test image with Pillow (in the venv) for the OCR/vision checks. - Objective grading (running generated code, exact-match answers) is deliberately used over LLM-as-judge to avoid judge bias.