# llm-bench A local LLM coding-model benchmark suite. Point it at any OpenAI-compatible endpoint, pick an instrument, and run. Pure stdlib Python — no packages needed. ## Quick start ```bash cp .env.example .env # set ENDPOINT + MODEL for your model python bench.py coding-suite validate # prove tasks are well-formed (no model needed) python bench.py coding-suite run # run all 24 tasks against the model ``` Results land in `results/` as JSONL + JSON. Work dirs in `work/` (both gitignored). ## Instruments ``` python bench.py [args...] ``` | Instrument | Command(s) | What it measures | |---|---|---| | `coding-suite` | `validate` / `run` / `report` | 24 single-turn bug-fix + tool-calling tasks, real toolchain (Go/Rust/Swift/TS) | | `agentic-fix` | `validate` / `run` | Multi-turn tool loop — find a cross-file bug in a small Go codebase | | `agentic-build` | `validate` / `run` | Build a Go+SQLite web app end-to-end (4 product gates: build/test/serve/persist) | | `agentic-multitask` | `validate ` / `run ` / `cohort ...` | Multi-task gauntlet: Go+HTMX, Rust CLI, Next.js+Prisma | | `router` | `validate` / `run` / `report` | Dispatch/router eval (26 ROUTE + 8 ORCH, deterministic scoring, privacy axis) | | `needle` | `run [--depths 1000,16000 --needles 3 --repeats 2]` | Needle-in-a-haystack retrieval + speed per depth | | `loop-battery` | `run` | CoT runaway probes (5 prompts × N runs, repetition scoring) | | `no-think` | `run` | Latency win from suppressing reasoning (think vs no-think A/B) | ### Examples ```bash # Coding suite (single-turn, ~5-15 min) python bench.py coding-suite run # Agentic build (multi-turn, ~10-40 min) python bench.py agentic-build run # Full gauntlet — all three real-app tasks, sequential python bench.py agentic-multitask cohort go_htmx rust_cli ts_next # Router/dispatch eval python bench.py router run # Needle retrieval at specific depths python bench.py needle run --depths 1000,16000,64000 # Override config per-run without editing .env ENDPOINT=http://10.0.0.40:8000/v1/chat/completions \ MODEL=qwen3.8-27b \ MAX_TOKENS=16000 \ python bench.py agentic-build run ``` ## Configuration Copy `.env.example` to `.env` and edit. Env vars win over `.env`, so you can override per-run (see the last example above). | Var | Default | What | |---|---|---| | `ENDPOINT` | `http://127.0.0.1:1234/v1/chat/completions` | OpenAI-compat chat URL | | `MODEL` | `local` | model id | | `API_KEY` | (blank) | bearer token (if endpoint needs auth) | | `MAX_TOKENS` | `8000` | per-turn generation cap | | `TEMPERATURE` | `0` | sampling temp (0 = greedy, recommended) | | `MAX_TURNS` | `40` | agentic loop cap | | `HTTP_TIMEOUT` | `900` | per-request wall timeout (s) | | `NOTHINK` | (blank) | `1` = suppress reasoning via empty think prefill | ### Compatible endpoints Anything that speaks OpenAI `/v1/chat/completions`: - **LM Studio** — `http://127.0.0.1:1234/v1/chat/completions` - **Ollama** (local) — `http://127.0.0.1:11434/v1/chat/completions` - **vLLM** / **llama-server** — `http://127.0.0.1:8000/v1/chat/completions` - **Ollama Cloud** — `https://api.ollama-cloud.com/v1/chat/completions` - **OpenAI** — `https://api.openai.com/v1/chat/completions` Tool-calling tasks use the OpenAI `tools` param. The XML/JSON tool-calling tasks use freeform text, so they work even on backends without native function-calling. ## Requirements - Python 3.10+ - Language toolchains for the tasks you want to run: | Toolchain | Needed for | |---|---| | Go 1.22+ | agentic-fix, agentic-build, go_htmx, go bug-fixes | | Rust 1.70+ | rust_cli, rust bug-fixes | | bun 1.0+ | ts bug-fixes | | Node.js 18+ + npm | ts_next | | Swift / Xcode | swift bug-fixes (macOS only; skipped gracefully elsewhere) | No Python packages needed — everything uses the stdlib. ## Adding your own tasks **Coding suite** — add a dict to `BUGFIX_T1` / `BUGFIX_T2` / `TOOLCALL` in `bench/tasks/coding_suite.py`. Bug-fix tasks need: `id`, `lang`, `symptom`, `buggy`, `fix`, `test`. `validate` proves buggy fails + fix passes. **Multitask gauntlet** — drop a `.py` file in `bench/tasks/multitask/` exposing `TASK_NAME`, `PLAN`, `STARTER`, `REFERENCE`, `validate(work_dir)`, `gates(work_dir)`. Auto-discovered. See `go_htmx.py` for the shape. **New instrument** — add a module under `bench/tasks/` with a `main()` function, register it in `INSTRUMENTS` in `bench.py`. Import from `bench.core` (`client`, `runners`, `tools`, `reporting`). ## Project layout ``` llm-bench/ ├── bench.py # CLI entry point ├── .env.example # config template ├── bench/ │ ├── core/ # client, runners, tools, reporting │ └── tasks/ # instruments (one per file) │ ├── coding_suite.py │ ├── agentic_fix.py │ ├── agentic_build.py │ ├── agentic_multitask.py │ ├── router.py │ ├── needle.py │ ├── loop_battery.py │ ├── no_think.py │ └── multitask/ # gauntlet task modules (auto-discovered) │ ├── go_htmx.py │ ├── rust_cli.py │ └── ts_next.py ├── results/ # generated (gitignored) └── work/ # generated (gitignored) ``` ## License MIT. See LICENSE.