llm-bench
A local LLM coding-model benchmark suite. Point it at any OpenAI-compatible endpoint, pick an instrument, and run. Pure stdlib Python — no packages needed.
Quick start
cp .env.example .env # set ENDPOINT + MODEL for your model
python bench.py coding-suite validate # prove tasks are well-formed (no model needed)
python bench.py coding-suite run # run all 24 tasks against the model
Results land in results/ as JSONL + JSON. Work dirs in work/ (both gitignored).
Instruments
python bench.py <instrument> <command> [args...]
| Instrument | Command(s) | What it measures |
|---|---|---|
| coding-suite | validate / run / report | 24 single-turn bug-fix + tool-calling tasks, real toolchain (Go/Rust/Swift/TS) |
| agentic-fix | validate / run | Multi-turn tool loop — find a cross-file bug in a small Go codebase |
| agentic-build | validate / run | Build a Go+SQLite web app end-to-end (4 product gates: build/test/serve/persist) |
| agentic-multitask | validate <task> / run <task> / cohort <t1> <t2> ... | Multi-task gauntlet: Go+HTMX, Rust CLI, Next.js+Prisma |
| router | validate / run / report | Dispatch/router eval (26 ROUTE + 8 ORCH, deterministic scoring, privacy axis) |
| needle | run [--depths 1000,16000 --needles 3 --repeats 2] | Needle-in-a-haystack retrieval + speed per depth |
| loop-battery | run | CoT runaway probes (5 prompts × N runs, repetition scoring) |
| no-think | run | Latency win from suppressing reasoning (think vs no-think A/B) |
Examples
# Coding suite (single-turn, ~5-15 min)
python bench.py coding-suite run
# Agentic build (multi-turn, ~10-40 min)
python bench.py agentic-build run
# Full gauntlet — all three real-app tasks, sequential
python bench.py agentic-multitask cohort go_htmx rust_cli ts_next
# Router/dispatch eval
python bench.py router run
# Needle retrieval at specific depths
python bench.py needle run --depths 1000,16000,64000
# Override config per-run without editing .env
ENDPOINT=http://10.0.0.40:8000/v1/chat/completions \
MODEL=qwen3.8-27b \
MAX_TOKENS=16000 \
python bench.py agentic-build run
Configuration
Copy .env.example to .env and edit. Env vars win over .env, so you
can override per-run (see the last example above).
| Var | Default | What |
|---|---|---|
| ENDPOINT | http://127.0.0.1:1234/v1/chat/completions | OpenAI-compat chat URL |
| MODEL | local | model id |
| API_KEY | (blank) | bearer token (if endpoint needs auth) |
| MAX_TOKENS | 8000 | per-turn generation cap |
| TEMPERATURE | 0 | sampling temp (0 = greedy, recommended) |
| MAX_TURNS | 40 | agentic loop cap |
| HTTP_TIMEOUT | 900 | per-request wall timeout (s) |
| NOTHINK | (blank) | 1 = suppress reasoning via empty think prefill |
Compatible endpoints
Anything that speaks OpenAI /v1/chat/completions:
- LM Studio —
http://127.0.0.1:1234/v1/chat/completions - Ollama (local) —
http://127.0.0.1:11434/v1/chat/completions - vLLM / llama-server —
http://127.0.0.1:8000/v1/chat/completions - Ollama Cloud —
https://api.ollama-cloud.com/v1/chat/completions - OpenAI —
https://api.openai.com/v1/chat/completions
Tool-calling tasks use the OpenAI tools param. The XML/JSON tool-calling
tasks use freeform text, so they work even on backends without native
function-calling.
Requirements
- Python 3.10+
- Language toolchains for the tasks you want to run:
| Toolchain | Needed for | |---|---| | Go 1.22+ | agentic-fix, agentic-build, go_htmx, go bug-fixes | | Rust 1.70+ | rust_cli, rust bug-fixes | | bun 1.0+ | ts bug-fixes | | Node.js 18+ + npm | ts_next | | Swift / Xcode | swift bug-fixes (macOS only; skipped gracefully elsewhere) |
No Python packages needed — everything uses the stdlib.
Adding your own tasks
Coding suite — add a dict to BUGFIX_T1 / BUGFIX_T2 / TOOLCALL in
bench/tasks/coding_suite.py. Bug-fix tasks need: id, lang, symptom,
buggy, fix, test. validate proves buggy fails + fix passes.
Multitask gauntlet — drop a .py file in bench/tasks/multitask/
exposing TASK_NAME, PLAN, STARTER, REFERENCE, validate(work_dir),
gates(work_dir). Auto-discovered. See go_htmx.py for the shape.
New instrument — add a module under bench/tasks/ with a main()
function, register it in INSTRUMENTS in bench.py. Import from
bench.core (client, runners, tools, reporting).
Project layout
llm-bench/
├── bench.py # CLI entry point
├── .env.example # config template
├── bench/
│ ├── core/ # client, runners, tools, reporting
│ └── tasks/ # instruments (one per file)
│ ├── coding_suite.py
│ ├── agentic_fix.py
│ ├── agentic_build.py
│ ├── agentic_multitask.py
│ ├── router.py
│ ├── needle.py
│ ├── loop_battery.py
│ ├── no_think.py
│ └── multitask/ # gauntlet task modules (auto-discovered)
│ ├── go_htmx.py
│ ├── rust_cli.py
│ └── ts_next.py
├── results/ # generated (gitignored)
└── work/ # generated (gitignored)
License
MIT. See LICENSE.